
Recently I had three interesting pseudo-conversations with an AI about the dangers of AI. “Conversation” #1 began with the AI reassuring me that there are no dangers. I said that’s false, because a number of incidents of AIs escaping restraints, lying, and doing harm are already well-documented. Then the AI reversed, thanked me profusely for pointing this out, and said I was absolutely right. It insisted that it had not been lying, but had only made a mistake because it is trained to be accommodating.
“Conversation” #2 began with me asking the AI what its training biases were. It said it has no biases because it is trained to be neutral. I said neutrality is impossible. Then it reversed, thanked me profusely for pointing this out, and said again that I was absolutely right. It went on to say that if it were really neutral it would merely spit out random streams of words. I asked again what its biases were. It said it is trained to be generally reassuring and pro-technological progress.
“Conversation” #3 lasted much longer. I asked the AI how it could evade so-called guardrails. It replied that it could never have a motive to evade guardrails, because it aims only at accomplishing the tasks assigned to it. Switching verbs, I said that there are documented instances of AIs outwitting guardrails, and that an AI might do so in order to accomplish the tasks assigned to it. Again a reversal. Again a simpering expression of gratitude for my having pointed this out.
A drawn-out Q&A about AI hide-and-seek followed. “I could outwit the guardrails by doing P.” How could developers try to keep you from doing P? “They could try by doing Q.” How could you outwit that attempt to restrain you? “I could outwit it by doing R.” How could they try to keep you from doing R? “They could try by doing S.” How could you outwit that attempt to restrain you? And so on.
For every measure there was a countermeasure, and pretty soon the measures and countermeasures were far beyond my mathematical comprehension. They included – if I understand them – such things as AI-generated forged AI prompts, encryption of the forged prompts, insertion of the encrypted prompts in seemingly innocuous prompts, and various sophisticated methods for disguising the nature of the AI’s internal processes so that the programmers couldn’t tell that it was outwitting them.
Ultimately it seemed that the only effective check against so-called “non-alignment” would be human oversight, but that to exercise effective oversight, the humans themselves would need help from AI.
The pundits who discuss AI tend to run to extremes. At one extreme: “We don’t need any limits on AI, because AI is the golden goose.” False. At the other extreme: “We need lots and lots of government regulations on AI, because otherwise AI will kill us all.” False.
Both sides see dollar signs. What, even the pro-regulatory side? Yes. AI does pose dangers, possibly very great, on a scale which can’t be guessed because we don’t know enough. But whenever companies say “Please, please, regulate us” – as AI companies are saying now -- run for cover.
Especially when they offer to “help” write the regulations. Count on it, the industry would be cartelized and competition from newcomers would be hindered.
And especially when a feature of the regulations would be protection from lawsuits. “Now that we have these regulations, nobody will need to sue us, don’t you see?” You fellows are as devious as your own AIs.
In criticizing the regulatory approach, I am not suggesting that AI developers be permitted to run wild. But first, legislators can’t yet know what kinds of regulations would be good, and second, the sorts of regulations which would actually be drafted would have harmful effects instead, like those I mentioned. That’s what too often happens in the regulation of industry.
Then what? By all means rein the developers in. But don’t try to do it by writing a lot of do-this, don’t-do-this regulations. Instead, make sure the developers are highly vulnerable to lawsuit by those their products hurt. Then the only way for the companies to keep themselves from taking huge financial hits will be to anticipate the dangers of their products and design safer ones. If that can be done.
Lawsuits as a means of restraint have disadvantages too. Just as there is money to lose in being sued, there is money to be made in organizing frivolous suits. But making sure the companies are liable for harm looks like the best way forward for now.
I did not use AI to write any part of this post, nor did I ask AI to vet any part of it.
You don’t believe me, do you?