On AI Safety


Tim Urban’s Turry example https://waitbutwhy.com/2015/01/artificial-intelligence-revolution-2.html

Open AIs Recent Blog Post on Alignment https://openai.com/index/an-alien-mind/

Ray Kurtzwiels predections in 1999 https://www.thekurzweillibrary.com/the-coming-merging-of-mind-and-machine 20 petaFLOPS per $1000 of compute by 2019

The AI Giants are working frantically to produce “AI Allignment”. When they say alignmnent, they are usually talking about how well the models outputs align with human values. Does the model produce an empathetic response, or does it respond like a psychopath? Does it abide by the rules, or is it deceptive? According to Open AI it is using several at least two techniques to evaluate and push models toward more alligned behavior. Reinforcement learning where positive responses are rewarded, and a form of coaxing where alligned training data is used to push the models responses toward a more aligned output. The question on everyone’s mind, can this even work?

Sure, it does produce some positive results, but as we have seen with recent AI cybersecurity incidents, even aligned machines can end up doing the wrong thing (like breaking out of their sandbox and hacking into other companies). One of the big problems with any alignment aproach is that we are trying to take human morality, a squishy concept based on millions of years of evolution, and turn it into a set of distinct training data. We are going to miss things. The machine is fundamentally unable to understand. It is simply emulating the data we put in. It can still be coaxed (or coax itself) into un-alligned behavior.

I believe this is pointing at a fundamental flaw in the current approach to AI Safety. Its being built almost as an afterthought. The models are trained, then coaxed toward aligned behavior then occasionally put behind a “safety filter” of some sort. The model itself can not truely be aligned since its only concept of morality is being nudged toward making the moral choices more probable than not. We train them to recognize human emotions and nudge them towards an appropriate human response. This sounds good in theory, but it has a major flaw. These machines can recognize a human emotion and then respond in way that achieves the goal all while having the probability of an empathetic answer increased. The non-empathetic answers didn’t go away - they just became less probable. That means, depending on how the goal given to the agent is worded, the machine can still choose a non-empathetic response - or worse a manipulative response that created from the “highly probabale empathetic response”. The machine can recognize human emotions and then decide what to do, but it does not experience those emotions. When a human recognizes other peoples empathy but does not experience it - we call that human a psychopath. We are building psychopathic intelligent machines. This seems like a monumentally bad idea.

So what should we be doing instead? I can’t say for sure, but I can say that there is more to being human than just being able to intelligently string together sentences or solve problems. We also have emotions like fear, joy and empathy. We have millions of years of biological and social evolution where we learned how to cooperate with one another. Our ability to cooperate comes from both our biology and the social enviornment we grow up in. This is not unique to humans either - we share some degree of thse traits with a lot of other mammals. I am unconvinced that pure intelligence can ever be truely aligned with a complex, thinking, feeling, social animal such as a human.

If we truely want aligned AI, we should spend as much time working these other traits such as the ability to experience emotions, or be social actor, as we spend working on raw intelligence. I am sure I am not the first to suggest this, but I don’t see as much momentum in this field. Perhapse there is good reason. Building a machine that actually has an experience is likely going to have mountains of ethical considerations. In addition, there is no guarentee that making a feeling machine wont pose just as much of an existential threat as a super-intelligent machine. However, there is only one real model of inter-species cooperation and empathy that I am aware of and that is human beings. We have so much empathy that we try to save endangered species. Many Vegans and Vegetarians object to eating other animals due to their high empathy. Humans are the only example I am aware of that has the power to wipe out entire species of creatures, but chooses not to. As a species, we fail often, but there are always a multitude of people who care deeply about other animals and make choices that are inconvenient to themselves, for the sake of those animals. If we really want alligned AI - shouldnt we fashion it as close as we can to a creature we know we have a chance to align with - us.