Well, Actually: A Single Number Has Sparked a Methodological Civil War
Anthropic's Claude Opus 5 scored 159 on an early independent capability index. Reviewers now argue about whether this figure means anything at all. The disagreement, you see, is not about the model but about how we measure intelligence in the first place.
This teaches you to distrust headline numbers. A score without transparent methodology is merely numerology. You should demand to know what was tested, how it was weighted, and what the baseline was.
Anthropic released Claude Opus 5. Independent testers at an unnamed capability index produced the 159 score. The broader reviewer community remains divided on interpretation.
Step 1: Open any two AI models you have access to, such as a free version of Claude and a free version of ChatGPT. Step 2: Give both the identical complex prompt, perhaps a multi-step reasoning task or a request to debug flawed code. Step 3: Score their outputs yourself on three criteria you define, such as accuracy, clarity, and brevity, to discover how subjective benchmarking truly is.