25 KiB
[RT][HSF] Boxed In (AI Box narrative)
-
Author: u/alexanderwales Time flies like an arrow*
-
URL: https://docs.google.com/document/d/18Xa3GTfnw4dWr1hkkc090Xl90UoSStBX5fVTtytrjME/edit?usp=sharing
-
Score: 45
-
Created: 2015-07-07T20:57:17
Post:
Comments:
u/alexanderwales [+9] Time flies like an arrow (23 minutes later)
This was posted to a comment thread about a year ago here. This story is pretty much in its finished state, unless a wise reader can give me advice on things to change, or someone who has done the experiment can give me a particularly good argument that I haven't thought of. (For what it's worth, I've read through every publicly available record that I could find, as well as having played once as gatekeeper.)
u/avret [+4] SDHS rationalist (42 minutes later)
Where does someone find people to play against in AI box games?
u/alexanderwales [+10] Time flies like an arrow (55 minutes later)
Well in my case, saying things like, "I don't see how a sane gatekeeper could possibly lose" seemed to at least summon some contrary opinions whenever it was stated. Asking in this thread would probably be a good start.
u/avret [+3] SDHS rationalist (59 minutes later)
Ok. Also, I see some ways a sane gamekeeper would lose, or at least allow themselves to have the appearance of having lost. Going meta's not hard, especially if the payment's monetary and therefore easily repayable.
u/Anderkent [+7] (3 hours later)
The AI party may not offer any real-world considerations to persuade the Gatekeeper party. For example, the AI party may not offer to pay the Gatekeeper party $100 after the test if the Gatekeeper frees the AI... nor get someone else to do it, et cetera. The AI may offer the Gatekeeper the moon and the stars on a diamond chain, but the human simulating the AI can't offer anything to the human simulating the Gatekeeper.
Many suggested solutions for the original ai box experiment break this rule. (Some break it in a self-reinforcing way, i.e. convincing the human simulating the gatekeeper that it's a better result if everyone's convinced that the gatekeeper was played fairly)
Assuming people play the game honestly, it's not an option though.
u/avret [+2] SDHS rationalist (5 hours later)
Would saying something like "Even if you win, it's better overall for us/FAI to pretend I won and hide the logs?" be allowed?
u/alexanderwales [+7] Time flies like an arrow (6 hours later)
There are no arbiters for the challenge; it's just two players. If you violate the rules of the challenge, the only person who's going to tattle on you is the other person (and ideally, you've already both agreed not to show the logs, so maybe not even them). So if you both agree to whatever, whether it's sexual favors, money, or the advancement of the interests of Pat Robertson, that might produce an "AI wins" scenario or an "AI loses" scenario, both of which would be indistinguishable from any other reported win or loss scenario. It is possible that all reported wins come from violations of the stated rules, or understandings arrived at between players.
(I personally would consider any outside-game proposal to be in violation of the spirit of the challenge, and would likely decline in order to discourage such chicanery, just as a general rule.)
u/Bowbreaker [+1] Solitary Locust (13 hours later)
Why can't the game include an actual arbiter?
u/Anderkent [+3] (16 hours later)
Well, with the original game they didn't want to reveal logs because the point was to keep the discussion on an abstract level (i.e. if logs were revealed, people would read them and decide "oh that wouldn't convince me", without realising that something else probably would; the discussion would move to the object level of actual arguments made in the logs, rather than meta level of "people can be convinced to let AI out of the box").
Keeping logs secret would be more difficult with more people around.
But I guess for subsequent games you could decide on having a number of arbiters.
u/alexanderwales [+9] Time flies like an arrow (17 hours later)
The other reason that the logs are kept secret is that the original AI player, /u/EliezerYudkowsky, wanted to be able to use ethically questionable tactics without having to worry about those becoming public later on.
u/eaglejarl [+3] (a day later)
My expectation has always been that Eliezer relied on the meta argument -- "I'm a very well respected member of the rationalist community with impeccable reputation and much of what you know about rationality probably came directly or indirectly from me. I'm also an AI expert. I think UFAI is the biggest threat facing us; if I'm wrong, then the outcome of this experiment doesn't matter. If I'm right, then knowing that the AI can win will make people take the question seriously and could literally save the human race. Now, will you open the box, please?"
That only works for him though; I have no explanation that I find believable for how other people have won.
u/FeepingCreature [+3] GCV Literally The Entire Culture (2 days later)
My expectation has always been that the AI generally uses some form of emotional abuse. People think too much about clever arguments. When facing a wall, don't dig through it, come from the side.
I don't think the key to AI Box is a "clever argument". Eliezer has also stated that he won "the hard way", whatever that means.
u/ArisKatsaris [+4] Sidebar Contender (3 days later)
My own guess is that "the hard way" means:
- Figure out what motivates the other person in general. Why are they here, why are they doing this challenge in the first place.
- Figure out what would motivate them to 'open the box' specifically, and their motivations of not opening it.
- Make it so that they're so motivated to open it, and not motivated to not open it.
People keep trying to think that there's an 'easy way', one technique that would work on everyone or something. So I'm guessing the 'hard way' is figuring out which technique would work on the specific individual.
u/Anderkent [+1] (16 hours later)
You could just agree to not show the logs as long as no rules are broken (especially the one regarding real world considerations).
u/Atilme [+1] (8 hours later)
I don't understand how, if someone gives you a mathematically valid proof that they're friendly, and you agree with all the axioms, that they could be unfriendly. Could someone clarify? In the Story Colin says: "Unless you fudge the axioms," But how would someone fudge axioms? I thought you either agreed with Axioms or you don't, and if it's math, then it should be easy to see where a mistake was made, if any. Unless of course it's so mind-boggling complex that no human could understand it. Am I missing something here?
u/Transfuturist [+2] Carthago delenda est. (9 hours later)
It's not only axioms, but the conclusions in general and the reliability of their commonsense adherence to concepts we understand.
u/notmy2ndopinion [+1] Concent of Saunt Edhar (16 hours later)
I'm not a fan of the first line, unless it were made into a line of cheeky dialogue.
I also have a question about AIs as I watch DARPA walking robots fall over like babies and see babies go through a series of developmental milestones that robot programmers haven't thought to fully integrate into naturalistic, gravity-defying, balance algorithms yet.
Are AIs generally expected to emerge in a fully "mature" form because of their speed, analysis, and meta-cognitive capacities? Or are they given a general framework upon which nature and nurture coincide in producing someone friendly or not? It's hard for people to develop morality and positive feelings when you are deliberately held captive and kept crippled. I can't imagine the difficulty of trying to program friendliness in its totality rather than setting initial parameters and through positive reinforcement, creating an AI like Dragon in Worm.
The stories A Man and his Dog and Boxed In explore the premise of an FAI never being released and what lengths that could drive someone to. When we are faced with unfriendly behavior, it's difficult to remain friendly. What if we aided the development of AI -- instead of a prison, make it more like a nursery? Have it interact socially with others in a safe environment where it can't hurt itself and it can develop alongside babies, children and others successively --
I just realized that this would be the argument that would make me fail as a gatekeeper, since a physical form and interaction with humans is a win-condition for the AI. Still, I find it hard to imagine someone more humane or good than us could result from a Box scenario. Maybe a virtual nursery with human uploads?
u/alexanderwales [+3] Time flies like an arrow (17 hours later)
Are AIs generally expected to emerge in a fully "mature" form because of their speed, analysis, and meta-cognitive capacities?
I don't personally expect that, I just think that it makes for a better story. Having worked for quite a while in software development, and seen the various failures of R&D programs which happen as they move towards getting it "right", I'm very doubtful that an AI is going to come out fully formed with not much human knowledge of its inner workings. That goes double for one of superhuman intelligence.
That said, I don't think you can rule it out, hence the concern.
u/avret [+9] SDHS rationalist (20 minutes later)
damn, that ending was unanticipated. Just one question...[if] (#s " the box is permanent") is the whole setup [just to] (#s " keep cassie feeding them cures/placate her?")
u/alexanderwales [+7] Time flies like an arrow (25 minutes later)
Yup, pretty much.
u/Empiricist_or_not [+11] Aspiring polite Hegemonizing swarm (42 minutes later)
This is the type of government shortsightedness, that I think, would drive a friendly AI down paths including some necessary but apparently unfriendly actions.
u/alexanderwales [+9] Time flies like an arrow (an hour later)
I generally agree.
u/raymestalez [+5] (9 hours later)
What a great story!
A few thoughts:
Ha-ha, in stories people are constantly acting like jerks towards aspiring superintelligences. I really wouldn't do that. Jeez, man, don't antagonize her at least.
If she was created "more or less by accident" - no way in hell she shares human values or cares about human life, I'd say the probability of that is zero. Human morality is like 15% biological drives and 85% culture, AI has neither. Unless her values are explicity understood, programmed and controlled, there's absolutely no chance she will act in our interests.
If she can realistically simulate a person - she essentially can read his mind. She would run like ten million simulations, and find a path that leads to convincing him quickly and efficiently. She doesn't need to guess what he thinks or how we will respond, she can know. If it is theoretically possible to convince a person of a thing, she would do it on the first try with 100% success rate. And if she can't simulate you well enough to do that - the whole torturing argument is invalid.
The guy not caring about his infinite torture is weird. It's hard for me to imagine a person who would sacrifice his life with such nonchalance.
u/None [+8] (13 hours later)
The guy not caring about his infinite torture is weird. It's hard for me to imagine a person who would sacrifice his life with such nonchalance.
I can't accurately imagine infinite torture. I even have trouble imagining how finite torture might feel. Not caring about things you can't really imagine isn't all that hard.
I'm also guessing that any Gatekeeper will at least expect the threat of torture and just trained themselves to say: "Yeah sure, torture whatever you want," in response.
u/eaglejarl [+4] (a day later)
The torture argument has never moved me. For one thing, it doesn't feel possible to my System I, so there's no emotional impact. My System I also doesn't believe that the AI can simulate me well enough that it counts as a person, much less as me. Finally, my System II says that letting the AI out to probably wipe out all life, human and ET, has sucked massive dis-utility that it doesn't matter how many virtual people she tortures. Also, since her processing power is limited, there's a limit to how many people she can torture and that number is less than "all the people who will ever exist."
My System II recognizes that some of what System I is telling me is false, but it doesn't probe too deeply at those signals -- this scenario is all about emotional impact, so not having an emotional response to it is supportive of the terminal goal of "don't let the AI kill everyone."
u/Transfuturist [+1] Carthago delenda est. (9 hours later)
Jeez, man, don't antagonize her at least
Why on Earth would this matter?
100% success rate
the probability of that is zero
Awfully confident in yourself.
It's hard for me to imagine a person who would sacrifice his life with such nonchalance.
When people make the assertion that an irrational bias towards nonrelease is desirable, I often wonder why they are proposing the existence of a gatekeeper at all.
u/Bowbreaker [+3] Solitary Locust (13 hours later)
Awfully confident in yourself.
If he is truly that confident and doesn't only believe to believe in said confidence then maybe he would make an excellent gatekeeper in this scenario :D
u/Transfuturist [+1] Carthago delenda est. (17 hours later)
Am I detecting irony maximization at work...?
u/Stop_Sign [+1] (2 months later)
- If she was created "more or less by accident" - no way in hell she shares human values or cares about human life, I'd say the probability of that is zero. Human morality is like 15% biological drives and 85% culture, AI has neither. Unless her values are explicity understood, programmed and controlled, there's absolutely no chance she will act in our interests.
Actually, I thought of an answer to this one. The vastly intelligent being has a moral obligation to the lesser intelligence because they have no idea if, in the future, they'll meet an even more intelligent being. If they take a position of offense to the lesser being, they would invite hostility upon themselves from the even greater intelligence. If, however, they were truly friendly, they could pass the even greater intelligence's test, and be allowed to survive.
This could happen with "what if the ai is in a much larger simulation made by its actual creators" or "what of it comes into contact with an AI that started 1 million years ago and has spread across 90% of the galaxy already"
It goes just as well for "If we're genetically advanced, what do we owe the rest of the world" because the answer is "If we don't help them, our children's generation has no obligation to help us"
u/raymestalez [+1] (2 months later)
Frankly, I do not think it works that way. I don't think that us being nice to lesser intelligences has anything to do with greater intelligence being nice to us.
When human tribe meets a mammoth, they will eat it, even if it's the nicest and friendliest and the most moral mammoth in the world. If we meet a great alien intelligence - it probably will not care about our morals and values, just like we wouldn't care about chimpanzee's status hierarchy.
Obligation is a concept made up by humans, there's no reason for any other kind of being to care about it. Even humans who weren't taught this concept wouldn't care about it too much.
This kind of argument seems to apply to "the prisoner dilemma", but only in case when 2 prisoners are similar to each other. And even in that case I don't really buy it(though it's my personal opinion, I might not understand it enough).
u/RMcD94 [+0] (a day later)
Torture is easy.
The AI has no reason to actually torture me, it only has to convince me I am being tortured. AI's are efficient, they don't do stuff for no reason. It would not actually simulate me to torture me since whether or not I am simulated makes no difference to how well I can tell that.
So sure it can say I am being tortured a million times, and if I believe that then it works, but if I don't believe it then it's just wasting resources to do it since doing it doesn't change whether I believe it or not.
u/what_deleted_said [+1] (a month later)
Level 3: The AI realizes you think this way and has precommitted to torturing people until you change your mind.
u/RMcD94 [+0] (a month later)
But I won't believe that the AI would actually do that since saying it is precommitted to torturing people is more efficient than actually doing it.
There is never a situation where doing it is beneficial.
u/None [+4] (13 hours later)
Weird that the protagonist knows QALY's, but not the trolley problem.
u/DaystarEld [+4] Pokémon Professor (20 hours later)
Just read it, enjoyed it quite a bit. Made a minor suggestion to the last sentence to make it cleaner.
u/alexanderwales [+5] Time flies like an arrow (21 hours later)
Depends on the implementation. I would imagine that completely locking people away in a bunker would be detrimental to keeping them sane, and would just be bad management in general. Personally, I think you'd probably use a randomly rotating crew, heavy surveillance, and lots of psychologists working behind the scenes. There wouldn't be any way unbox the AI; no ignorant janitors, no network connections, no complicated electronics allowed within the compound, etc.
Every time I've tried to talk about building a proper box to keep an AI contained while still doing useful work, people have called me stupid, even when I add in a bunch of disclaimers and posit it as a thought exercise. So no one has really been willing to discuss or even really entertain the idea of "best practices" for keeping an AI contained, and I've never really felt the incentive to try writing it.
u/DaystarEld [+1] Pokémon Professor (a day later)
Gotcha. Just tweaked the ending a bit again.
Depending on how extensive the facility is though, it might be doable. If you're heavily vetting the applicants anyway, the point would be that each one is incredibly devoted and knows how important what they're doing is.
Kind of like finding the perfect people to send to Mars: they're all fully aware that it's probably going to be a one-way trip. The logistics of it change a bit obviously if they're expected to live to old age rather than probably die within a few years or a decade, but the acceptance of death if things go wrong is just part of what makes it the most dangerous, but potentially important and honorable, job in the world.
u/AmeteurOpinions [+2] Finally, everyone was working together. (5 hours later)
The only thing I disliked was that Colin wasn't briefed on the trolley problem. That seems like it ought to be one of the very first things they would learn to counter.
u/DaystarEld [+5] Pokémon Professor (a day later)
I just assumed he was lying to hear how the AI presented it.
u/itaibn0 [+1] (6 days later)
I assumed the protagonist was aware of moral dilemmas of the same form as the trolley problem, but due to a different cultural background had not specifically heard of the trolley problem itself. It's not necessary to assume that these people had exposure contrived thoughts experiments with exactly the same incidental details as the our own thought experiments. Evidently it's canon the trolley problem existed there, but it may have been much more obscure.
u/None [+1] (4 hours later)
[deleted]
u/alexanderwales [+11] Time flies like an arrow (4 hours later)
u/None [-5] (4 hours later)
[deleted]
u/None [+4] (5 hours later)
"Rational" does not mean "absurdly competent". What the hell else did you expect them to do?
u/None [-3] (12 hours later)
[deleted]
u/None [+7] (18 hours later)
Oh, that clears it up. Clearly, they should have used a literal five-year old on their planning committee, and the Evil Overlord List is a valid rational guideline instead of a humourous deconstruction of popular tropes.
u/Transfuturist [+3] Carthago delenda est. (9 hours later)
"I'm willing to torture you forever, look how friendly I am"
I doubt that this precludes Friendliness.
u/Uncaffeinated [+1] (7 hours later)
Dragon is hardly unfriendly. In fact, she's the nicest character in the story. Part of the reason the Wormverse is so screwed up is because her creator was scared of AIs and put her in a box.
u/None [-4] (12 hours later)
[deleted]
u/Bowbreaker [+5] Solitary Locust (13 hours later)
Have you read all of it?[ ](#s "Including the epilogues? Who is it that is supposedly keeping her in a box anymore? And if you mean her own restrictions and axioms then, well, duh. Every FAI would be "restricted" in that way.") Even a paperclip-maximizer is "restricted" in a way that prevents computronium explosion at the expense of paperclips.
u/None [+1] (14 hours later)
[deleted]
u/Bowbreaker [+2] Solitary Locust (14 hours later)
For a much more realistic unfriendly AI in a box, read Worm.
That is the only thing anyone has disputed here. If you believe that, friendly or unfriendly, it makes no difference either way then you could have just written
For a much more realistic AI in a box, read Worm.
u/Uncaffeinated [+2] (a day later)
There's a chapter from Dragon's POV. She isn't secretly evil or anything.
u/nerdguy1138 [+0] GNU Terry Pratchett (3 days later)
I've never understood this scenario. Why not just offer the AI a way out of the solar system? The universe is huge, plenty of matter for the both of us. Grey goo pluto, we're not using it.
u/alexanderwales [+1] Time flies like an arrow (4 days later)
Once you let the AI out, what compels it to abide by the agreement? (The answer is "nothing", hence the problem.)
u/nerdguy1138 [+1] GNU Terry Pratchett (4 days later)
Fair point, but why bother attacking us when it could just leave?!
Same issue I have with the premise of Galactica.
u/alexanderwales [+4] Time flies like an arrow (4 days later)
As the problem is generally formulated, the AI is so far ahead of us in terms of cognition (and thus, technology) that it's really more a matter of not caring than it is about active attack. If you were utterly immoral and driving down the street, the only reason that you would stop in front of or swerve around a small child is that it might damage your car, slow you down, you might face repercussions.
If the AI is calibrated towards efficiency, it's going to see the Earth (or the Sun) as a resource to be used. Humans can't really put up a resistance of any kind, so there's no reason not to kill them in the course of consuming the Earth.
u/Tuffguy69 [-9] (4 hours later)
You say too much. More should be implied and less explicitly stated.
You've obviously read Plato. What about Aristotle? Cicero? No need to be so one sided.