CJ Trowbridge / Projects

AGI still isn't real.

A critical look at why today's artificial intelligence systems still do not constitute artificial general intelligence.

ARC-AGI-3 chart for Astra

The concept of AGI is the idea that there is some moment when AI systems are equivalent in capability to humans. The problem with this idea is that there is no one-dimensional measure of intelligence or capability (either in humans or in AI systems). This is not a real term in research. It’s a marketing term used by people who never studied computer science or artificial intelligence, or by hype-bois with AI psychosis who are jsut regurgitating the slop being fed to them by marketers, confindently resassured by sycophantic LLMs that their latest ChatGPT session is somehow going to, “usher humanity into a new era of harmonic convergence.”

I know a lot of smart people and I know a lot of stupid people. I also know a lot about the capabilities of AI systems. There is no AI that reaches the level of someone like Andrej Karpathy. Conversely, there are many humans I wouldn’t trust for a second if any chatbot told me they were wrong. This idea that there is some line there is just insane; it’s just obviously wrong. The concept of AGI never made sense to any thoughtful person who really worked on AI. Sam Altman himself has recently started to concede this point about AGI, saying, “it’s not a super useful term.

Nevertheless, OpenAI’s president Greg Brockman claimed yesterday that this new model they are about to release means we have entered, “the AGI era.” Note that he didn’t claim the new model is AGI or that they have solved AGI. He dissembled that it’s up to the reader to decide. Predictably, all the headlines ran with his dubious, indirect claim that this model is an example of AGI.

OpenAI can’t decide what definition they are using for AGI.

Sam Altman has claimed at various points that AGI is when AI can do the work of, “a median human coworker,” or when it can, “match or exceed humans at most economically valuable tasks,” or when it can do, “a significant amount of the work in the world..”

Brockman cites a single benchmark score; ARC-AGI-3.

Brockman tweeted a partial screenshot of the ARC-AGI-3 chart for Astra, claiming 'arc-agi-3 is now saturated.'

Several things immediately jump out to me from these numbers; mainly the things Brockman’s screenshot crops out from the report.

ARC-AGI-3 chart for Astra

First of all, OpenAI spent $20-$50 thousand dollars in tokens for each of these results. This benchmark is primarily simple visual puzzle problems. And the humans whose score sets the 100% mark on the chart were paid $115 per 90-minute session. There is no indication of how long these $50,000 Astra sessions took, but it’s difficult to imagine a world where anyone would make a serious argument that the net present value of $50,000 in tokens is equivalent to $115 for 90 minutes of human labor, especially on a problem like solving simple visual puzzles.

ARC-AGI-3 puzzle problems

Secondly, the two separate series shown on the chart are the model’s 17-63% scores, and then the model’s much higher scores with some unspecified tools being used. So the score they’re talking about is not the model itself, it’s the tools.

So the bottom line for me is that regardless of how you try to spin these results, this is not a model that is anywhere near the level of a human. (Even if you’re willing to spend $50k per prompt.)

The AGI Narrative Serves Another Purpose

In my graduate thesis, The Illusion of Understanding: Deconstructing AI Metaphors, I argue that AI-industry language repeatedly and deliberately borrows human concepts such as intelligence, learning, attention, rewards, and punishment, then uses them to tell inaccurate stories about what’s going on, not in service of understanding, but in service of marketing and for investors’ sake.

These inaccurate metaphors do not make the technology easier to understand. They replace an explanation of what’s actually going on with meaningless anthropomorphic hand-waving. It’s pure chicanery. Brockman’s real goal in selling this dissimulating narrative is to increase the IPO price by hyping inaccurate metaphors about the capabilities of a chat bot we’ve all rightly come to view with suspicion and distrust.

This is not AGI.

It is not even close.


← Back to Projects

← Back to Home