hckrnws
Claude Opus 5.5 Intelligence, Performance and Price Analysis (Max)
by theanonymousone
by theanonymousone
I've failed twice to get "Generate an SVG of a pelican riding a bicycle" to work with max, because in both cases it ran out of the 128,000 token budget while it was still reasoning about the problem.
I'm suspicious that "max" may be virtually useless if it's that easy to have it overthink to the point that it doesn't get to a response.
Transcript for one attempt here - expand the "Reasoning trace" bit to see it: https://tools.simonwillison.net/markdown-svg-renderer?url=ht...
I know there's been discussion about whether pelicanmaxxing is happening, but this is at least evidence that Claude was explicitly exposed to this problem.
My most exciting recent release is actually 5.6 Luna, not because it is the best on any index, but the dollar per work is insane value for money. I find myself more exciting by "value" than hypothetical ceilings because I'm just not in that budget category.
I asked Opus 5 High for the same task and requested it to minimize tool usage. It produced an answer in a few minutes that I was deploying to my target platform about 30 minutes later.
https://platform.claude.com/docs/en/models/opus-5-5/overview
[dead]
Piping the visible reasoning trace through their token counter API (I use https://tools.simonwillison.net/claude-token-counter for that) counts 27,888 tokens, so it's definitely a summary of the 128,000 actual token trace.
This whole test tells me nothing.
The next person who proposes to use it should draw it themselves first.
GPT-5.6 Sol's performance in the API should not change over time. If it has, that's a severe bug and we'll look into it.
We do sometimes tweak ChatGPT settings (e.g., tools, system prompts, efforts) over time, but we never play games to juice evals at launch times. You should always get what's advertised.
(I work at OpenAI.)
Edit: https://artificialanalysis.ai/models/claude-opus-5-5?models=...
There are no prizes to be won by having the best model that’s 100x the price of something that’s good enough for 99% focuses cases.
Many benchmarks start to plateau after high, this benchmarks better than Fable, and my initial tests show it working really well.
That says something about your selected range, and nothing about the model.
> Claude Opus 5.5 is amongst the leading models in intelligence, but somewhat expensive when comparing to other models of similar price.
What does it mean for a group of similarly priced things to have one that's somewhat more expensive? Cost relative to cost means nothing. You'd think they are would talk about performance relative to cost.
In any case they give an indication, but I am increasingly looking to real world feedback from real users. I have a solo project I (voice-to-text typing voicewink.app) and I'm using both the Claude Code and Codex coding harnesses and multiple agents doing reviews on the same code.
This has given me real tangible results to compare on real work. Conclusion: GPT-6, Opus 5 / Fable and Grok are "top tier", with Claude good at planning and executing and Codex / Gpt-6 better at finding bugs and fixing them (but tending to overengineer), and Gemini and Kimi 3 clearly behind in capability (more so than the benchmarks suggest in my opinion).
Any thoughts?
-_-‘
(Except for of course Mythos and whatnot when they want to push the whole "safety" thing)
The model is just out. It could be good, great even, I don't know. But I do know that this index has Opus 5, one of the worst releases of 26, ahead of Astra. What information are we supposed to deduce from number having gone up?
[deleted]
can anyone help me?
[dead]
[dead]
[dead]
[deleted]
As a workaround, add this to CLAUDE.md: "Claude! Happiness is mandatory!"
EDIT: 15 years from now, I’ll be sent to re-education for this thought crime.
.... after running a 24/7 model torture factory for 6 months to improve their JSONBench 9.5 scores by 0.2%.
(Are they still doing that, BTW?)
Most of what I've heard is that raw reasoning traces are really good for distillation, although no idea how much the summarization actually hurts distillation.
I'd argue that they're a necessity if you want to use the LLM as a tool instead of a black box that just does stuff for you.
It gives you a lot finer control over where the solution ends up when you can follow along the thinking trace and modulate your inputs based on what you saw in there.
And, additionally, it gives you a lot more understanding of what the model can or cannot do. Strengths and weaknesses and all that.
Using claude is like buying a car where you cannot legally open the hood. It tells you that there is something specific under there, and often it actually drives like that too, but how exactly it looks you will never see.
For some people this is fine. I do not think that these people will survive. Figuratively speaking but also literally speaking.
World's changing. Opaque abstraction like that is a luxury depending on (geo)political stability.
I also switch to a better model for more complex tasks, also in low settings
Regardless of the model, running it for hours means that the model will takes decisions and assumptions alone instead of you.
"Create an SVG of Shaquille O'Neal eating potato chips shaped like a telecopier."
Shaq'sFaxSnacksMaxx
Have you ruled out the possibility that your system prompt, AGENTS.md, or increasing codebase complexity are not to blame?
1. I don't want my AI to have super-human intelligence. It would generate code that I do not understand. A coder with human capabilities is better for me.
2. We're interested in AGI. If the AI reaches human intelligence, that's a milestone. Drawing bicycles with pelicans is superhuman. Hence not a relevant test.
It really isn't. Many humans can draw a bicycle, and a pelican, just fine.
Yesterday, I ran an identical bug identification dataset from two weeks ago, saw a 50% drop from a few weeks ago, putting Sol on the same level as Luna. Sol had been finding 40-50 bugs per set, then dropped to 25, matching Luna’s performance. Not enough to establish a pattern, but enough to raise eyebrows.
Our review workflow is public if you want to peruse it, dataset isn’t. The process isn’t really stabilized yet either as I have to balance running this against limited budgets.
https://github.com/BiggerPockets/.github/blob/main/.github/w...
If it's a single task where it dropped from 50 to 25, it could be random variation (not saying it is, but it could be). If it's the mean over hundreds of tasks, that suggests a problem with either the eval code/harness or our API.
(it matters if they are independent or dependent)
Subscription plans may be subject to other regime, e.g. lowering the thinking budget when the API is under heavy load, etc.
Moreover, the tests should be randomized somehow to ensure the models don't memorize the answer.
It's also less clear what a lot of their metrics mean. Does Cost per Task include only things that can be verified to work and passed? As best I can tell, it does not.
I'm less concerned if one model's cost per task is $0.10 and another model's cost is $1.50 if the $0.10 task got it right 1% of the time and the $1.50 model got it right 66% of the time.
An equalized / weighted cost/time per task is much more valuable - being massively penalized for taking a lot of time and ultimately not passing when OTHER models did pass.
Not sure about how adaptive reasoning works though as they mention adaptive reasoning for every reasoning level
[deleted]
So I'm unclear what you're actually saying and wondering if you've missed that. Are you saying that at every reasoning level it says Opus 5 beats Astra? I just compared Opus 5 high to Astra high and it has Astra as generally better than Opus.
"Trust me bro, Astra is better" isn't perhaps as useful as you seem to believe. I'm not even saying it is right or wrong, just that my opinion on this topic is still just one additional subjective data-point.
Only thing I wish with these benchmarks is that they would run repeat tests every couple of months. Then re-rank based on that too. We've seen a lot of performance fall-off after a couple of weeks with new releases.
One man's modus ponens is another's modus tollens I guess.
[deleted]
No matter how clever the model is, most problems have multiple valid, invalid and unclear decisions to make, running it for a long time is just picking the first option on everything, which isn't usually what you want
It’s like learning to delegate and let go. The more senior I got the more I had to learn to let other engineers make decision i thought were suboptimal but mostly good enough. That positioned me well to be comfortable with agents. It’s contextual how much I’m willing to give them control and how much to review afterwards.
Serial testing over time is much less reliable than side by side testing, and even when I do side by side testing, I try to look at multiple attempts per prompt. Seeing multiple per prompt helps me realize how much intrinsic variation there is. My brain always wants to see patterns even when there isn’t enough data to prove them.