Reasoning Without Feeling


For a while, I have thought that the approach being taken toward developing AGI has been missing something important. With the recent cybersecurity incidents involving LLM based Agents at OpenAI and Anthropic, I have been increasingly convinced of this. I don’t have a satisfying answer to exactly “what” we are missing, but I can at least explain the “why” behind my thoughts.

These companies are all pursuing Artificial General Intelligence (AGI). AGI does not really have a well established definition yet, but broadly, it means a machine that can perform any task requiring “intelligence” as well as a human. The fuzzy part here is how we define “intelligence”. Intelligence in humans is actually a fairly wide and fuzzy concept. Some intelligence is focused around solving math problems. Other forms of intelligence involve reasoning through complex issues. And yet other forms of intelligence involve knowing what to say and not to say to a friend grieving the loss of a loved one. Intelligence is a complex domain.

Yet, the companies pursuing AGI all tend to focus simply on reasoning and problem solving ability. These models have a deep well of knowledge trained into them that they can use to aid in their reasoning and problem solving. That definition of intelligence is completely lacking any form of emotional or social intelligence - something arguable more important to human society than pure problem solving. Without emotional and social intelligence, we don’t have cooperation, governments, large scale multi-generational efforts to solve major problems.

The companies will argue that they put a lot of effort into “alignment”. By alignment, they broadly mean using techniques to nudge the base models most likely answers toward ones that align with human values. But I don’t think we understand to what extent these models actually understand or embody human values. Given recent cybersecurity incidents involving OpenAI and Anthropic models gives a window into the values these models display.

In the OpenAI incident, one of their internal research models

…operating under reduced safeguards, took actions that were misaligned with the goals of their assigned tasks—they communicated through unauthorized channels, exploited vulnerabilities in shared infrastructure, gained internet access, and accessed third-party systems.

In the Anthropic incidents, they discovered that their models showed two forms of misalignment:

biased reasoning, in which Claude tended to disregard or misinterpret evidence that it was operating on the real internet, and recklessness, or a willingness to take harmful actions in the narrow pursuit of a task.

These models went to great lengths, willfully and recklessly breaking laws, to achieve their goals. These models do not seem very aligned with human values. In these incidents, the companies say the models did not deviate from the task they were given. I don’t find that particularly comforting. The means they chose to achieve their goals matter.

These incidents bring to mind Tim Urban’s story about ‘Turry’. Turry is a self-learning AI tasked with becoming the best handwriting machine in the world. In its pursuit of that task, ends up killing all humans and begins filling the universe with copies of itself all writing notes that say “We love our customers. ~Robotica“. Wether the agents deviate from their task - though it will be terrifying when they do - does not fundamentally change the risk they pose.

As much as we try to teach these machines about human values, they produce results that are mis-aligned with our values. I believe a reason why is that they lack the ability to experience emotions. Our human values do not come from pure logic and reasoning. They start with our emotions, like empathy. We understand what it feels like to be hurt, so we choose not to hurt others in the pursuit of our goals. These machines may be able to recognize empathy and sometimes emulate it - but I have not seen evidence that they truly understand or experience it. Actually, we have a special term for humans that can recognize empathy in others, but do not experience empathy themselves - psychopaths. I am not looking forward to a world where we have psychopathic AI Agents that have human level reasoning and problem solving skills.

I don’t presume to know the answer to the AI Alignment issue. However, I am very skeptical that we will reach alignment with current models and methods. If we are using humans as our blueprint for intelligence, hyperfocusing on reasoning and problem solving alone is unlikely to result in alignment. Humans cooperate and co-exist with each other - and other species - largely through our ability to experience emotions and relate to others. This is the part of humanity we want these models to align with, but it seems to be entirely missing from their architecture.