The people building AI are the worst at predicting it

Quick answer

Eighteen dated AI predictions from 2024 and 2025, graded against their own words on their own deadlines. Capability forecasts mostly came true, some early. Deployment forecasts, agents joining the workforce, a billion Agentforce agents, programmers replaced within a year, failed almost uniformly. The top scorer was the skeptic Gary Marcus at A-, and being right earned him nothing.

Nobody goes back to check

In March 2025, Dario Amodei told the Council on Foreign Relations that AI would be writing 90% of code within three to six months, and essentially all of it within twelve. That twelve month deadline passed in March. Almost nobody went back to check.

That is the strange etiquette of AI forecasting. The predictions are loud, dated, and delivered with total confidence, and then the date arrives and everyone has moved on to the next one. There is no scoreboard. So I built one.

In February 2025, I scoped agent workflows into Whizi on the assumption Altman was right about 2025. I killed the project six weeks later, after doing the math on what a single multistep agent run costs against a flat monthly subscription. Six weeks of my own roadmap, spent on someone else's forecast. I pulled every dated 2024 and 2025 prediction I could find with a primary source and a deadline that has already passed. Eighteen made the cut.

The grading rule, so you can argue with it: a prediction is graded against what its own words promised, on its own deadline. If you said 90% by September and September brought 30%, "well, eventually" is not a grade. It is an excuse.

The scoreboard

Who, whenThe predictionWhat happenedGrade
Sam Altman, January 2025AI agents "join the workforce" in 202595% of enterprise pilots showed no P&L impact; the best agent finished 30.3% of office tasksC-
Marc Benioff, September 2024One billion Agentforce agents by end of 2025About 29,000 deals and $1B in annual recurring revenueF
Klarna, February 2024AI assistant doing the work of 700 agentsRehiring humans from May 2025 for lower qualityC
Dario Amodei, March 202590% of code within three to six months25% to 41% industry wide, 75% at Google by mid 2026C+
Dario Amodei, March 2025Essentially all code within twelve monthsNot closeF
Eric Schmidt, April 2025The vast majority of programmers replaced within a yearProgrammers still employed; entry level roles down about 16%F
Mark Zuckerberg, January 2025An AI midlevel engineer, the leading billion user assistant, and Llama as the top model, all in 2025None of the three; record pay packages for human engineers insteadF
Elon Musk, April 2024AI smarter than any human by end of 2025Grok 4 scored 25.4% on Humanity's Last ExamF
Mira Murati, June 2024PhD level intelligence "for specific tasks" in about 18 monthsA certified gold at the 2025 International Math OlympiadB-
Gary Marcus, January 202525 dated predictions, four graded hereAll four correct, and none of the year's successes on the listA-

The workforce that never clocked in

Sam Altman, January 2025: "We believe that, in 2025, we may see the first AI agents 'join the workforce' and materially change the output of companies."

MIT's NANDA project found in August 2025 that 95% of enterprise generative AI pilots produced no measurable P&L impact. Carnegie Mellon staffed a fake company entirely with AI agents and measured the office work they finished: the best model completed 30.3% of tasks. A 30% employee does not join your workforce. They exit your probation period.

One carve out saves this from an F: inside software companies, coding agents genuinely did change output. Grade: C-.

Marc Benioff, September 2024: "Our vision is bold: to empower one billion agents with Agentforce by the end of 2025."

Salesforce reported roughly 29,000 cumulative Agentforce deals by early 2026 and crossed $1B in annual recurring revenue. A real business, and about 34,000 times short of the promise, counting deals rather than agents. My favorite detail: Salesforce now reports "agentic work units served" (3.8 billion of them) instead of a count of agents. When you cannot hit the number, change the unit. Grade: F.

Klarna, February 2024: its AI assistant "is doing the equivalent work of 700 full-time agents."

Then in May 2025 CEO Sebastian Siemiatkowski reversed course and started rehiring humans: "We focused too much on efficiency and cost. The result was lower quality." The AI kept the routine tickets. The humans came back for everything that mattered. Credit for running the experiment in public and admitting the result. Grade: C.

The code that actually got written

Amodei's 90%. The direction was right and the denominator was wrong. Google says 75% of its new code is AI generated as of mid 2026, up from 25% in late 2024, though that counts every accepted autocomplete and humans still review before deploy. Industry wide estimates in early 2026 clustered around 25% to 41%.

So: 90% of all code in six months, no. Essentially all code in twelve, not close. A historic shift in how code gets written, absolutely yes. The prediction described a real revolution and got every number wrong. Grade: C+ for the 90%, F for "essentially all."

Eric Schmidt, April 2025: "in the next one year, the vast majority of programmers will be replaced by AI programmers."

The year is up. Programmers remain employed almost everywhere they were employed before. The real labor damage is narrower and crueler than the prediction: Stanford and ADP payroll data show employment for 22 to 25 year olds in the most AI exposed jobs down roughly 16% since late 2022 and still shrinking. Entry level took the hit; the vast majority kept their jobs and got Copilot licenses. Grade: F.

And the number nobody predicted at all: Cursor went from $100M to $4B in annual recurring revenue in about seventeen months, and on August 14 SpaceX closed its $60B acquisition of the company, the largest startup acquisition ever recorded. Not one 2024 prediction list I checked has "a code editor becomes the biggest startup exit in history" on it.

Meta's very expensive year

Mark Zuckerberg, January 2025, on Joe Rogan: "Probably in 2025, we at Meta... are going to have an AI that can effectively be a sort of midlevel engineer that you have at your company that can write code." On the earnings call weeks later he added two more for 2025: Meta AI as the leading billion user assistant, and Llama as the most advanced and widely used model.

None of the three happened. Llama 4 landed with a thud while Gemini 3 took the frontier crown in November 2025. And Meta spent the year making the opposite bet with its wallet: $14.3B for half of Scale AI, researcher pay packages reported between $100M and $450M, then 600 people cut from the superintelligence lab in October and roughly 8,000 layoffs in May 2026.

Meta predicted an AI that codes like a midlevel engineer, and instead spent eighteen months setting the world record market price for human ones. Grade: F.

Elon Musk, April 2024: AI "smarter than any one human probably around the end of next year."

By end of 2025, the flagship xAI model was Grok 4, launched as "the smartest AI in the world," scoring 25.4% on a benchmark literally named Humanity's Last Exam. The claim did not survive contact with the exam. Grade: F.

Mira Murati, June 2024: PhD level intelligence "for specific tasks" in about eighteen months.

For math it basically happened: Google DeepMind's Deep Think took an officially certified gold at the 2025 International Math Olympiad, years ahead of most 2024 forecasts. The qualifier she used, "for specific tasks," is doing heavy lifting, and it is exactly the qualifier her louder peers refused to use. The hedged version came true. The unqualified one, shipped as GPT-5's "PhD-level" branding in August 2025, went badly enough that Altman admitted OpenAI "totally screwed up" the rollout. Grade: B-.

The skeptic's report card

Gary Marcus published 25 dated predictions on January 1, 2025. The industry mostly rolls its eyes at him. Grading his gradeable claims:

  • "We will not see artificial general intelligence this year." Correct.
  • Agents will be "endlessly hyped throughout 2025 but far from reliable." Correct, see the entire first section.
  • "Profits from AI models will continue to be modest or nonexistent." Revenue exploded, but he said profits, and the labs are still burning cash. Correct on the letter.
  • Possibly no "GPT-5 level" consensus leap in 2025. GPT-5 shipped; the leap did not. Correct in spirit.

Grade: A-. The docked minus is for what the list does not contain: nothing in it predicted a $4B coding tool category, an IMO gold, or agents becoming genuinely useful in the one domain where output is checkable. The best forecaster of 2025 called every failure and missed every success.

And the uncomfortable second order fact: being right earned him approximately nothing. The investors who bet against his entire worldview are the ones who priced Cursor at $60B. Forecast accuracy and financial returns run on separate scoreboards, and only one of them compounds.

Capability predictions came true. Deployment predictions did not

Sort the 18 predictions into two piles and the noise disappears.

Predictions about capability, what models would be able to do, aged well and sometimes came in early. The length of tasks agents can finish has been doubling roughly every four to seven months, per METR, and the trend held.

Predictions about deployment, what organizations would let AI actually do, failed almost uniformly. Workforce agents, billion agent platforms, replaced programmers, AI employees: all of it hit the same wall of quality bars, liability, integration cost, and the stubborn fact that a system completing 30% of tasks is a demo, not a hire.

Altman said agents would join the workforce. They joined the IDE, because the IDE is the one workplace where a wrong answer gets caught by a compiler instead of a customer.

So here is the decision rule I use now, with my own money on the line: when you hear a capability prediction, take it seriously, even the wild ones, because the trend lines keep winning. When you hear a deployment prediction with a date attached, double the timeline, then check what the person saying it is selling. The same rule is how I pick which model to pay for: the benchmark is a capability claim, the workflow is a deployment claim.

Every quarter from here I grade whatever came due, and the next edition opens with my own dated predictions, so you can run me through the same rubric. Grading is free. Being graded is the price.

Workflow checklist
  • Grade a prediction against its own words, on its own deadline; "eventually" is not a grade
  • Take capability predictions seriously, even the wild ones; the trend lines keep winning
  • Double any deployment timeline that carries a date
  • Check what the person making the prediction is selling
  • Write your own predictions down with dates so someone can grade you back
Common questions

Frequently asked questions

Was Dario Amodei right that AI would write 90% of code?

Not on his deadline. He said 90% within three to six months of March 2025 and essentially all code within twelve. Industry wide estimates in early 2026 clustered around 25% to 41%, with Google at 75% of new code by mid 2026 counting every accepted autocomplete. The direction was right and the numbers were wrong.

Did AI agents join the workforce in 2025?

Not outside software. MIT's NANDA project found 95% of enterprise generative AI pilots produced no measurable P&L impact, and Carnegie Mellon's agent company test saw the best model finish 30.3% of office tasks. Coding agents inside software companies were the one real exception.

Who predicted 2025 in AI most accurately?

Gary Marcus, the skeptic, on the claims graded here: no AGI, agents hyped but unreliable, profits still modest, no consensus GPT-5 leap. All four held. His list also missed every success of the year, from the $4B coding tool category to the IMO gold.