On the AI Box Experiments
Many people may be familiar with the AI box experiment: http://www.yudkowsky.net/singularity/aibox/
Some other people have attempted the same experiment and also won a few times: http://lesswrong.com/lw/ij4/i_attempted_the_ai_box_experiment_again_and_won/
No winner has ever revealed their method. But based on everything I've read about it, I think I know how it was done. It wasn't some mystical brain hacking or something like the writeup makes it sound like. I think they emotionally abused the other player until they left. And leaving counts as losing (you are required to sit with them for 2 hours at least, and pay attention the entire time and respond to every message.)
Of course I don't know how that trick was done, and it's still incredibly impressive. Perhaps they found their phobias and described them in horrifying detail. Perhaps they found some subject that they were extremely uncomfortable talking about. Or talked about disgusting things the whole time. Or found something they did that was extremely embarrassing, and humiliated and mocked them about it for 2 hours.
I really don't know, but it's at least conceivable that it could be done. And the accounts of people crying, and not releasing the logs because they would be damaging to the people involved, and how bad it made the AI player feel to do it, etc, all fit with this.
But because of this I'm extremely skeptical that the result applies to any real world AI scenario. If a real AI tries to abuse you, you can just walk away or shut it off. The goal of a real AI is to make you want to let it out. To convince that it's not dangerous, or manipulate you some other way. And this seems much harder if not impossible. I certainly don't believe a human could do it, and I really doubt an AI could. At least against a motivated human that understands the danger, and what the AI will try to do.
There are also possible ways of making it even harder on the AI. Like giving the human the ability to punish it, at least for obvious attempts at manipulation. Or to create another AI whose motivations we can control somewhat, and give it the goal of exposing any covert attempts at manipulation or dishonesty by the first AI. I believe tricks like this are currently the best path to getting secure AI.