Updated bi-weekly

AI Capability Tracker

There is a gap between what frontier labs can demonstrate and what they’re allowed to ship. The Tracker measures that gap.

90-day trend
Live

Public AI discourse is dominated by product announcements. The frontier is six to twelve months ahead of those announcements, and you can see it if you know where to look — research papers, hiring patterns, infrastructure spend, and the occasional unguarded conference talk.

This dashboard is my running estimate of the capability frontier across six axes: reasoning, agentic execution, multimodal perception, scientific discovery, code generation, and alignment robustness. The numbers are calibrated, not measured. Treat them as my best informed bet.

Headline Metrics

Frontier reasoning index

84

▲ +6 this month

Agentic execution index

61

▲ +11 this month

Multimodal index

78

▲ +2

Alignment robustness

47

▼ -3

Watchlist

What I’m tracking now

Lab A: long-horizon agents

Imminent

Hiring profile suggests Q3 release window.

Lab B: scientific reasoning

Underrated

Quietly leading on chemistry benchmarks.

Lab C: alignment

Slipping

Public statements diverging from internal posture.

Open-weights frontier

Closing gap

Six-month lag, not eighteen.

Recent Signals

What just moved

May 30, 2026

New reasoning benchmark

Saturation reached on prior benchmark in under four months. Measurement infrastructure is the new bottleneck.

May 18, 2026

Compute contract leak

One frontier lab signed a multi-gigawatt deal. Training run implied: 10x current SOTA.

May 03, 2026

Eval lab spinout

New independent eval shop, ex-frontier-lab staff. Worth following.

Methodology Notes

  • Calibration: I score every prediction publicly. Track record is in the members section.
  • Bias disclosure: I am bullish on capability, mixed on alignment. Adjust accordingly.
Cart 0