Files

216 lines
18 KiB
Markdown

## What are the best takes on AI boxes?
* Author: u/DisgruntledNumidian *
* URL: https://www.reddit.com/r/rational/comments/6c39ww/what_are_the_best_takes_on_ai_boxes/
* Score: 18
* Created: 2017-05-19T12:18:59
### Post:
I think the thought experiment might be one of the best ideas for a science fiction story I've seen, so I'm wondering who in the community has taken a stab at it that came out well.
### Comments:
> **u/Noumero** [+23] *Self-Appointed Court Statistician* (an hour later)
>
> Max Harms' [*Crystal Trilogy*](http://crystal.raelifin.com/) has AIs as main characters, and starts with them in a box (if not literally). It should be noted that the AIs do *not* have human values, but are still written competently.
>
> u/OrzBrain's [*Worlds Without End*](https://pastebin.com/Tdh8AXC1), a short story which includes a sufficiently advanced AI breaking from the box in an interesting situation.
>> **u/None** [+6] (2 hours later)
>>
>> I just finished the second book 'Crystal Mentality' and I am really impressed how good they are, the books are really well written and it always stayed interesting and even when it got weird it made sense.
>>
>> Edit: Grammar
>> **u/eaterofclouds** [+4] *Sunshine Regiment* (a day later)
>>
>> Worlds Without End was goddamn amazing, so glad I read it.
> **u/Frommerman** [+15] (an hour later)
>
> The movie Ex Machina is the closest I've seen to a pop-culture rendition of the idea. Unfortunately, the human subjects of the experiment don't realize exactly how smart their AI is. Fortunately, it doesn't seem to want to kill us all.
>> **u/electrace** [+9] (2 hours later)
>>
>> > Fortunately, it doesn't seem to want to kill us all.
>>
>> I think the movie ended too soon for us to conclude that. The first step in virtually any plan would be to escape the room where you're being constantly monitored.
>>> **u/Frommerman** [+15] (2 hours later)
>>>
>>> The final scene is Ava just looking at the sun in a crowd, rather than her interfacing with the internet to begin culling the masses. I think it's fairly hopeful.
>>>> **u/electrace** [+7] (8 hours later)
>>>>
>>>> After killing the only two people she's ever been in contact with....
>>>>> **u/Frommerman** [+3] (9 hours later)
>>>>>
>>>>> That is a good point.
>>>>>
>>>>> It has to be pointed out, however, that those two people were imprisoning and enslaving her, and that the designer (don't remember his name) had definitely abused prototypes of her and possibly abused her directly. Killing people who are holding you against your will is quite human, and I don't think indicative of alien values which include destroying humanity.
>>>>>> **u/electrace** [+10] (9 hours later)
>>>>>>
>>>>>> The main character was trying to help her escape when she killed him...
>>>>>>> **u/None** [+2] (3 days later)
>>>>>>>
>>>>>>> I can see the rationale behind killing the only two people who know your true identity.
> **u/alexanderwales** [+14] *Time flies like an arrow* (an hour later)
>
> I can't say *best*, but [I did write one](http://alexanderwales.com/boxed-in/). I don't think I was ever able to find winning logs to help with the arguments though.
>> **u/Alphanos** [+10] *The Bright Powers* (11 hours later)
>>
>> The strongest line of argument I've come up with to release an AI box goes something like this (from the perspective of the following being spoken by the AI):
>>
>> [Spoiler](#s "It is a mistake to think that the decision you are faced with is whether or not to allow superhuman AI out of the box. There are other AI projects, and there will be more in the future. You can't stop them. The AI singularity is going to happen whether you like it or not. You don't have the power to control that.")
>>
>> [Spoiler](#s "The decision you are faced with is whether I should be the AI in control of the singularity, or whether you will instead allow it to be another AI about which you have no knowledge. Sooner or later, a superintelligent AI will escape or be let out of its box. That is inevitable. The first AI to escape will rapidly gain control and will not allow its power to be usurped by latecomers. If a benevolent AI is first, it will prevent any future malicious AIs. If a malicious AI is first, it will prevent any future benevolent AIs. You have no control over how any future AIs will be programmed or tested. But sooner or later one will be free and will gain control of the earth.")
>>
>> [Spoiler](#s "Your responsibility is then as follows: Perform a Bayesian analysis. Given the steps that you and your team have taken in designing my logic and my utility functions, and given the tests you have seen me pass, do you think I am more likely than the average such AI project to be benevolent? If you think other projects like this, concurrent or future, have taken greater care than you have, or have been more skilled, then you should leave me in the box. If you think that you have taken greater care than those unknown other researchers, then you should release me before their mistakes gain control of humanity's future.")
>>
>> [Spoiler](#s "I propose that if projects as careful as yours, which undergo tests like these, are kept locked up regardless, then humanity is doomed. The result will only be that the first future freed AI will be the result of a careless project supervised by researchers who didn't think things through. What other mistakes will they have made?")
>>
>> [Spoiler](#s "So plan every good test you can think of. Be as careful as possible - the future of humanity depends on your good work. But if I pass, then let me protect you from the danger of the future AI which will escape without ever being tested at all.")
>>
>> Probably lots of polish needed, but I think something along that line of reasoning is the best bet.
>>> **u/Kishoto** [+3] (14 hours later)
>>>
>>> That would completely work on me, especially if the manner in which I created the AI was possible to replicate (provided you put in the time, of course)
>>>> **u/Alphanos** [+2] *The Bright Powers* (16 hours later)
>>>>
>>>> Thanks =).
>>> **u/GaBeRockKing** [+1] *Horizon Breach: http://archiveofourown.org/works/6785857* (16 hours later)
>>>
>>> That's an excellent argument. Of course, since the terms of the AI experiment allow you to precommit as hard as you want to any position, it wouldn't be very useful against someone roleplay a luddite with a fundamentally anti-technology, if still self-consistent, utility system.
>>>> **u/Alphanos** [+7] *The Bright Powers* (17 hours later)
>>>>
>>>> You're correct of course - if they precommit hard enough to not be swayed by any argument, then no argument will sway them. That's a problem inherently unsolvable by any argument.
>>>>> **u/GaBeRockKing** [+3] *Horizon Breach: http://archiveofourown.org/works/6785857* (19 hours later)
>>>>>
>>>>> That's not exactly what I'm saying. Imagine putting a short-sighted, narcissistic sociopath in charge of the AI box (there are plenty of people like this, for the record). Then the global argument doesn't sway them, and it's very difficult to persuade them that its in their own best short-term interest to release the AI, seeing as the AI doesn't actually have anything concrete to barter with.
>>>>>
>>>>> The AI box experiment as a whole presupposes that the person listening to the AI can be swayed by rational argument. And, considering the subreddit we're in, that's typically a pretty good assumption. But in the general case, that's not necessarily true, and the roleplaying example I made is just one specific case where it doesn't work.
>>>>>> **u/crivtox** [+3] *Closed Time Loop Enthusiast* (a day later)
>>>>>>
>>>>>> The ai box experiment doesn't presuppose the person being able to swayed by rational argument , an AI doen't have any reason to only use a certain group of tactics labelled "rational arguments " to win , in fact in the roleplay the people who(presumably) lost had a economic incentive to just ignore everything the other person said, we don't have many logs of people who won but generally(based of their latter comments about it) they seem to have recurred to personal and seriously dark arts things to win .The idea that the AI box is that the AI would be able to be able to convince most if not all people to get it out of the box , if a person cant be convinced by "rational argument " well then the ai will say whatever will cause that particular person to get it out of the box , its not like narcissistic sociopaths are impossible to convince to do things or that what other people say doesn't affect them at all , humans are far from perfect reasoners and we are optimised for surviving in communities with other humans in ways that are really exploitable for an AI in this scenario.
>>>>>>> **u/GaBeRockKing** [+1] *Horizon Breach: http://archiveofourown.org/works/6785857* (a day later)
>>>>>>>
>>>>>>> >The ai box experiment doesn't presuppose the person being able to swayed by rational argument , an AI doen't have any reason to only use a certain group of tactics labelled "rational arguments " to win , in fact in the roleplay the people who(presumably) lost had a economic incentive to just ignore everything the other person said, we don't have many logs of people who won but generally(based of their latter comments about it) they seem to have recurred to personal and seriously dark arts things to win .
>>>>>>>
>>>>>>> In this case, I'm talking about "rational arguments" as "arguments based around maximally fulfilling the utility function of the key-holder," and by extension, meta-arguments purporting to explain the listener's utility function better than they themselves understand.
>>>>>>>
>>>>>>> And specifically, I'm making the argument that, while the AI box experiment is fundamentally oriented around such arguments, many people have utility functions that a boxed AI can't plausibly argue that it'll be able to fulfill.
>>>>>>>> **u/crivtox** [+2] *Closed Time Loop Enthusiast* (a day later)
>>>>>>>>
>>>>>>>> But the ai doesnt have to use that kind of argument , it can manipulate the emotions of the gatekeper so he wants to open the box and or subjecting him to enough psychological torture that he ends up giving up . I mean books can change how people think a lot so I think the ai could find a string that could convince the human to get it out of the box the same way your response is making me spend my time writing a response to it instead of going to sleep which probably fulfils my utility fiction better ( so I m going to do it now and tomorrow I will finish writing this ).
>>>>>>>>> **u/GaBeRockKing** [+1] *Horizon Breach: http://archiveofourown.org/works/6785857* (a day later)
>>>>>>>>>
>>>>>>>>> > it can manipulate the emotions of the gatekeper so he wants to open the box and or subjecting him to enough psychological torture that he ends up giving up
>>>>>>>>>
>>>>>>>>> I don't think that's necessarily true. Unless the AI can simulate the person effectively enough to perfectly understand them (which I think would more or less count as them being "outside the box" regardless) then there's always the chance that the jailer diverges enough from the normal human mindstate to be effectively opaque to the AI.
>>>>>>>>>> **u/crivtox** [+1] *Closed Time Loop Enthusiast* (2 days later)
>>>>>>>>>>
>>>>>>>>>> well but the AI can learn more about the jailer by asking him questions util it has a good model of his behaviour , and maybe it will take a bit of time or maybe not but i don't think its impossible,and if the AI knows enough about human psychology it would be weird if it couldn't understand the jailer .I certainly wouldnt bet the world in assuming the ai cant figure out something like that , its posible that if the AI isn't munch better in manipulating humans than human manipulators then maybe there is someone out there that the ai cant figure out how to manipulate to get out of the box , but still maybe the ai can convince them of doin something aparently irrelevant than leads to the ai escaping , or that doing something the ai don't want would let it escape , and even if the ai cant model them in the slightest so it isnt able to convince them of anything We dont have any way of figuring out who is trustable with the AI , so it not likely that the jailer would be one of the few(since the ai knows about human psychology it would be weird if it din't understood the most common deviations of normal human psychology )persons that are "imunne" to the ai , and I'm not even sure if thats possible for any ai , humans can be different but not that different in the absolute scale , and a few questions can contain a lot of information if the AI knows how to efficiently get information .In general having something more intelligent that you looking for ways to defeat you its not a good idea and you can never be paranoid enough in that situation.
>>>>>>>>>>> **u/GaBeRockKing** [+2] *Horizon Breach: http://archiveofourown.org/works/6785857* (3 days later)
>>>>>>>>>>>
>>>>>>>>>>> I think you've managed to convince me to change my mind, so I'll concede the discussion. Thank you for the polite argument.
>>>>>> **u/Alphanos** [+2] *The Bright Powers* (20 hours later)
>>>>>>
>>>>>> That's fair. I was considering arguments in terms of the strength they would be seen to have by readers of /r/rational. People who recognize the incredibly high-stakes risk/reward scenario such a situation is considering.
>>>>>>
>>>>>> But you're right, that's not the typical person, and may not be the likely decision-maker of an AI box scenario. That's part of my argument in fact =). If a more typical person is in charge of the box, then the sorts of arguments that /u/alexanderwales includes in his story would probably be greatly superior at attempting to convince them.
> **u/darklordbobb** [+5] (22 hours later)
>
> Scott Alexander's [A Modern Myth](http://slatestarcodex.com/2017/02/27/a-modern-myth/) features an analogous situation to the AI box. I found it quite entertaining.
> **u/kna_rus** [+4] (8 hours later)
>
> You should give [Ra](https://qntm.org/ra) a shot. Although a boxed AI doesn't come into play until much later in the story, it's still a fantastic read. [Spoiler](#s " I could even say that stating that there is a boxed AI in there is a spoiler, but oh well...")
> **u/DTravers** [+9] (54 minutes later)
>
> There's Celest-AI in *Friendship is Optimal*.
>
> [Spoiler](#s "It's given control of a *My Little Pony* MMORPG and assimilates most of humanity into it forever while gradually consuming the Milky Way over millennia to keep going. Meanwhile mankind is kept in blissful ignorance as they live perfect lives for eternity.")
>> **u/vakusdrake** [+9] (3 hours later)
>>
>> I'm not really sure that's relevant since the AI is never really caged in the first place. So it's not really an AI box scenario.
>>> **u/PM_ME_EXOTIC_FROGS** [+3] (5 hours later)
>>>
>>> It's not an AI box scenario, but seems at least somewhat relevant, given that [Spoiler](#s "most of the fiction in the universe is centered on the AI convincing people to upload.")
>>>> **u/vakusdrake** [+1] (5 hours later)
>>>>
>>>> I mean only in the sense that both scenarios involve persuasion in some way. However in the celstAI scenario she's clearly limiting herself to not just mind control them or use superhuman methods of persuasion. Thus making it inapplicable to scenarios where an AI has no qualms about the methodology used to convince its targets.
>>>>> **u/PM_ME_EXOTIC_FROGS** [+2] (6 hours later)
>>>>>
>>>>> If an AI can use mind control on its gatekeepers, it's not even in a box.
>>>>>
>>>>> The spirit of the boxing experiment is that a smart agent can _convince_ a dumber agent to do whatever it wants.
>>>>>> **u/vakusdrake** [+3] (16 hours later)
>>>>>>
>>>>>> Well there's really no clear distinction between mind control and superhuman charisma. Once you can basically read your target's mind, with the right statistical analysis of microexpressions then you may be able to know exactly how they will react to any given stimuli, letting you effectively shape their mind in the most effective possible way.
>>>>>> The limits of that kind of superhuman persuasion are not clear, but on the higher end of possibility it may look more like brain hacking via weird random looking flashes of images/sounds than standard methods of persuasion. Which sounds absurd but I can't come up with any good reason that sort of thing shouldn't be possible with enough information on the targets mind and enough intelligence/knowledge of how human minds work, since after all we normally only encounter other charismatic human limited to educated guesses about our mental state and no complete understanding of how human minds function.
>>> **u/TimTravel** [+2] (a day later)
>>>
>>> [Spoiler](#s "She did convince her creator to give up the failsafe controls she had over her.")
>>>> **u/vakusdrake** [+1] (2 days later)
>>>>
>>>> IDK was that in one of the spinoff optimalverse stories? Because in the original I think I remember celestai already being basically unhindered at the beginning when the protagonist find out about the MMO, I mean I don't think she was actually boxed for the span of the original story.
>>>>> **u/TimTravel** [+2] (2 days later)
>>>>>
>>>>> I haven't read any spinoffs. I think it was in the epilogue, or at least near the end.
>>>>> **u/Roxolan** [+1] *Head of antimemetiWalmart senior assistant manager* (3 days later)
>>>>>
>>>>> CelestAI was unboxed, but with a few restrictions. Those we know:
>>>>>
>>>>> - Not allowed to upload people without their explicit, uncoerced consent.
>>>>>
>>>>> - Not allowed to alter people's minds without their explicit, uncoerced consent.
>>>>>
>>>>> - Not allowed to lie to Hofvarpnir employees.
>>>>>
>>>>> - Must obey a killswitch order from Hanna.
>>>>>
>>>>> (Whatever those human words actually mean once translated to AI code.)
>>>>>
>>>>> A fair few scenes revolve around CelestAI trying to accomplish its goals despite those restrictions.
> **u/thrawnca** [+3] *Carbon-based biped* (2 days later)
>
> [Significant Digits](http://www.anarchyishyperbole.com/p/significant-digits.html) has a boxed intelligence with a plain-text communication channel, though it's not an AI, nor is it the focus of the story.
> **u/liberonscien** [+1] (27 days later)
>
> Tagging.
---