Jack Clark (co-founder of Anthropic) published an analytical essay on the current state of the AI industry.
His main thesis: with a probability >60%, we will see a fully automated R&D cycle by the end of 2028. That is, a top-tier model will be able to autonomously train its next, smarter version without the involvement of meat bags.
Sounds like the setup for a cheap sci-fi, but Clark explains everything with metrics. If you look at the aggregated benchmark data, it becomes a bit unsettling how quickly the horizon of human competence is collapsing.
1️⃣ Saturation of engineering benchmarks
SWE-Bench (solving real GitHub issues). At the end of 2023, the top result was a measly 2% from Claude 2. Now Claude Mythos Preview hits 93.9%. In fact, the benchmark is passed.
The autonomy timeframe (METR) has grown from 4 minutes for GPT-4 in 2023 to 12 hours of continuous work for Opus 4.6 today. That's a full working day for a mid-level developer who doesn't go out for smoke breaks. By the end of 2026, a jump to 100 hours is expected. That's enough for a neural network not just to fix a bug, but to rewrite the service architecture while you're busy with something else.
2️⃣ Automation of data scientist routine
MLE-Bench benchmark (offline Kaggle competitions). The best system based on Gemini3 hits 64.4%.
CORE-Bench (reproducing scientific ML articles from repositories). The model itself installs dependencies, sets up the environment, runs code, and checks results. Score — 95.5%.
3️⃣ Most important: AI ventures into hardcore optimization
Neural networks have learned to write and optimize GPU kernels (Triton/CUDA). Anthropic measures how models optimize code for training LLMs on CPUs. Over a year, the speedup increased from 2.9x (Opus 4) to 52x (Claude Mythos Preview). A human would need a full working day for such a result.
PostTrainBench has emerged, where large models fine-tune small open-source weights. Currently, AI agents already achieve ~50% of the performance of top ML engineers from leading labs.
Yes, models still lack "creativity" to create new paradigms (like transformers), but they don't need it. Extensive scaling through autonomous teams of AI agents will be more than enough.
It may well come to the point where large companies are "large" in terms of capital (need lots of GPUs) but very small in terms of people. Why keep a staff of 50 engineers if one architect can manage a swarm of 500 synthetic agents that don't burn out and write code at the speed of light? Well, that's if everything goes according to the ideal scenario and there are no black swans.
Do you believe?
Komentar
0Belum ada komentar.
Masuk untuk ikut berdiskusi.