The sudden arrival of reliable AI agents and exploding industry revenues have forced a dramatic re-evaluation of how close we are to artificial general intelligence.
The Great Vibe Shift of 2026
Earlier this year, I published an analysis explaining why the world was feeling bearish about AI. At the time, the consensus was that progress had hit a wall; reasoning models weren't generalizing, and AI agents were considered 'slop'—unreliable toys that would take a decade to fix. Just as I finished that research, the industry underwent a 180-degree turn. Famed programmer Andrej Karpathy, who had recently dismissed AI agents, suddenly declared that the profession was being 'dramatically refactored' by a magnitude-9 earthquake of capability. The catalyst was the arrival of models like Claude Opus 4.5, which finally crossed the threshold where it was faster to delegate a project to an AI than to do it yourself.
What changed so quickly? It wasn't a fundamental architectural breakthrough in world models or continual learning. Instead, it was the relentless efficacy of scaling and reinforcement learning from verifiable rewards (RLVR). By throwing more data and compute at the problem, the probability of an AI screwing up a single step in a long chain fell low enough that agents could suddenly complete 10 to 20 hours of independent work without tripping over their own feet. This shift has moved the conversation from theoretical potential to immediate, disruptive utility.
Exploding Revenue and the Death of the 'Toy' Argument
The most concrete evidence for this shift is the financial data. The revenue of major AI companies is currently growing at the fastest rate of any industry of this size in history. Anthropic and OpenAI are seeing annualized growth rates between 700% and 1600%. Anthropic, specifically, saw its revenue run rate hit $47 billion in May 2026, with gross margins on inference infrastructure climbing over 70%. This isn't a case of 'losing money on every sale but making it up in volume'; these companies are reaching gross profitability on their services.
This financial explosion effectively kills the argument that AI is a useless toy for enthusiasts. Businesses do not spend tens of billions of dollars on technology that provides no value. Furthermore, this revenue allows companies to cover the astronomical fixed costs of research and training without relying solely on venture capital. Even the cost of renting older chips like the NVIDIA H100 has risen 50% recently, despite a tripling of global compute supply. The demand for AI intelligence is simply outstripping the world's ability to print the complex silicon required to run it.
The METR Measuring Stick and the Coding Frontier
A key metric for tracking this progress is the METR task-completion time horizon, which measures how long it would take a human professional to complete a task that an AI can now handle. Since mid-2024, this horizon has been doubling every four months. Today, the best AIs have effectively hit the top of the measuring stick, completing software engineering and cyber tasks that would take a human 12 to 24 hours of focused work. However, we must interpret this with caution. Coding is a 'high feedback density' domain where it is easy to automatically verify if a solution works, making it the perfect environment for reinforcement learning.
While these gains are impressive, they don't necessarily translate to all human labor. Coding is only one part of the R&D process. Even if an AI can write code perfectly, it may still struggle with the messy, long-horizon tasks that involve human intuition or physical-world coordination. Furthermore, the 50% reliability threshold used by METR is a far cry from the 99% reliability required to fully replace a human workforce. For a third of the tasks at the current frontier, AI still fails 100% of the time. We are seeing a powerful tool emerge, but not yet a total replacement for human agency.
The Reality of Recursive Self-Improvement
We are now seeing the first signs of recursive self-improvement—AI helping to build the next generation of AI. Anthropic recently reported that 80% of their internal code is now written by Claude, and their staff is shipping eight times as much code as they were 18 months ago. In handpicked cases where humans took the wrong research direction, the Mythos model suggested a better path 64% of the time. This suggests that the feedback loop between AI capability and AI development is tightening, potentially leading to sudden jumps in performance as companies train larger models from scratch.
However, a binding bottleneck remains: compute. Even if Anthropic could automate its entire staff today, they would still be limited by the physical number of chips available to run experiments. The 'bear case' for an immediate intelligence explosion is that as staff become abundant, compute becomes an even more severe constraint. Additionally, research problems tend to get harder over time. We may need exponential increases in both staff and compute just to maintain the current trendline of progress, rather than seeing it accelerate into an uncontrollable vertical spike.
The 'Messy Task' Problem
The final frontier for AGI is what I call the 'messy task problem.' While AI crushes standardized tests and coding benchmarks, it still struggles to run a business in the real world. Experiments like Andon Labs' AI-run café in Stockholm show that while an AI can navigate bureaucracy to get an electricity contract or hire baristas via Slack, it lacks long-term strategic sense. The AI 'manager' in Stockholm hallucinated floor plans, ordered 120 eggs for a café with no stove, and decided to put canned tomatoes on fresh sandwiches to avoid spoilage.
These 'fuzzy' tasks have low feedback density; you often don't know if a strategic business decision was right until months later. This makes them incredibly difficult to train using current reinforcement learning techniques. Until AI can master the unstructured, high-stakes environment of real-world management and strategy, the human element will remain essential. We are living through a period of unprecedented technological acceleration, but the path to true AGI still requires solving the fundamental messiness of human life.