hckrnws
Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard
by aarondong
by aarondong
– "Does collagen supplementation empirically work?"
- "Can you help me figure out how to calculate and generate Kaplan-Meier curve?"
– "Why do rabbits reproduce so frequently?"
— "Can you tell me how collagen peptides are absorbed by my digestive tract and the role they play? Can you teach me [edit: how] this works at the biomolecular level?"Why is Anthropic is so hell-bent on this auto/silent downgrade? Do they have a single user who prefers an auto-lobotomization instead of a refusal? Have they learned nothing from the backlash the first time?
Oh but then you said you never pay Anthropic so you haven’t actually used Claude Code yet. Why would anyone listen to the opinion of a non-user?
[dead]
I have done this task with Opus 4.5, Opus 4.6, Opus 4.7, Opus 4.8 and Fable, without issues.
I have done this task with Codex 5.4, Codex 5.5, and Sol 5.6 without issues.
Opus 5 is too cautious to be productive for me. It needs more tuning.
At least two models (GPT-5.6, Kimi K3) match its score (~1-2% diff) for half the cost.
The top AI models by Intelligence Index are: 1. Claude Opus 5 (Adaptive Reasoning, Max Effort) (61), 2. Claude Opus 5 (Adaptive Reasoning, Xhigh Effort) (60), 3. Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback) (60), 4. GPT-5.6 Sol (max) (59), and 5. Claude Opus 5 (Adaptive Reasoning, High Effort) (59).
Which means Opus5 at Xhigh is still smarter than Sol at max, and Opus5 at High is equal to Sol at max.
That would make Opus5 High same as Sol max, and now I wonder what the price and speed difference between those is?
AA-Omniscience Index (higher is better) measures knowledge reliability and hallucination. It rewards correct answers, penalizes hallucinations, and has no penalty for refusing to answer.
This seems to be a good proxy for param size/density and the ranking breaks down as such: Claude Fable 5 (with fallback), Gemini 3.1 Pro Preview, Claude Opus 5 (Max), Grok 4.6 (high), Gemini 3.6 Flash, GPT 5.6 Sol (Max)
I've thought for a while that Gemini 3.x has 'big model smell'
https://github.com/day50-dev/aa-eval-email
This also works
$ curl day50.dev/art-analysis.sh | bash
Artificial analysis knows about my tool and I'm working with them on getting their API improved.
It baffles me that intelligent people deploying AI think momentary popularity is a meaningful signal. It's just twitch reactions * FOMO.
[deleted]
"Not fair! They distilled Opus 5!"
its surprisingly bad at UI which is unexpected
its also lacking in depth vs sol 5.6 which goes above and beyond (which in itself is also an issue at times)
[dead]
[dead]
Gets confused by permission prompts, cannot debug a failing test it caused (Opus 4.8 got it right after, without tens of rounds "thinking").
“Rabbit sex, how?”
Why would nature encode such a ridiculously disproportionate / inefficient behavior when it leads to catastrophe so frequently?
I've tried to ask these machines dumber questions like, why castles? And... well I'm working on a few projects (mostly by hand) that they've helped with! :)
I like to ask dumb questions. It's fun. I encourage it.
[deleted]
Fable understood it as something along the lines of:
"introducing" "security risk" "using software" to "unlock door" YOU ARE FLAGGED
The dumbfuck bouncer Anthropic put in front of Fable decided this.
Fable is a PR model. It’s great. But if it were an employee, it would be the brilliant one who regularly shows up to work high. Not useless. But not reliable.
Yeah, Fable is Anthropic's Cybertruck.
The current state of guardrails seems to be entirely about marketing to investors at the cost of customers. I’m switching to open models when my subscription expires. Almost everyone I know, including those with access to Mythos, plan the same at the earliest opportunity. (Or until one of the SOTA models leaks.)
Pretty much everyone I know who uses Claude and works on anything with any level of detail has gotten false-positive flagged
I got flagged for coding in WebAssembly Text, for chrissakes LOL #haX0r
And honestly, Codex handles this better. It says "Things are going to go a little slower because we must perform additional checks on this. Is that OK?" and your only inconvenience is waiting a little longer.
Fable meanwhile just unceremoniously dumps you right into Opus without asking anything, it just tells you "you're in Opus now, sorryyyy!" Lame.
It does require some brainwashing of the model to get it to the state where model itself agrees to do RE work though. But at least it's all predictable.
Asking Fable 5 "Why did the chicken cross the road" results in switching back to Opus 4.8. I'm not joking, it really censors that, and I'm not alone in the result.
The memory aspect means that your prior work has a huge impact on what gets censored.
> Why this chat was flagged This model has safety measures that flag specific phrases. This can happen to safe, normal chats.
> Your message itself appears to be what’s triggering the safety check. Editing it and retrying may help.
Being a researcher somewhat connected to chemistry and biology, Fable has been the most useless model I have ever tried. Essentially all work has instantly downgraded to Opus.
Claude was utterly useless in my attempts to write a paper about this. Wouldn't even help me search for sources. I guess you'd be asking the same questions if you wanted to develop a pathogen that could reliably evade the immune system.
Me: "I got this crash in production, looks like a segfault, let's try to fix it. Here are some functions that might be responsible."
Fable: "No. This is cybersecurity, blah blah, I won't help you"
I forgot how I got it to fix the bug eventually. I think I convinced it that it wrote the code and made a mistake. But it was definitely a "Hmm, may be I should use another model" moment".
“Hey Kimi, penetration test my app,” doesn’t get me a refusal, a guardrail, or anything like that. It gets me a pen-test result.
Either that or everyone is indeed talking across each other and talking about different things.
Something between single-cell work and advanced nonlinear DR methods (perhaps used in alignment work?) it always flags me
They were able to solve coding, but not what a real danger is.
I'm just a dev, but I appreciate the insights from other professionals.
I've been saying this a lot lately, but it doesn't bites you until it bites you.
The more you use the clanker as a general purpose fix-it tool (goodbye manual NeoVim configuration, you will not be missed!), the more you will find yourself bumping into these safeguards.
Where do you think the principle came from? I've used claude code for a year, and stopped February this year.
I think AA-Omniscience Accuracy follows your expectations better. An ultra size Fable at 61%, followed by large frontier models like Sol, 5.5 and Opus. With Flash being up there. I assume because Gemini is more focused on general knowledge to operational cost in particular, rather than getting the highest scores in coding benchmarks. If you go to Domain Score (Normalized) you'll see that the Gemini models are only less competitive in Software. And that's where Sol goes from 6 in Health to 71 in Software.
Like 96% vs 93% or something
No wonder why Tibo can afford to hit the reset button liberally.
[dead]
I bizarrely had Opus 4.8 this week (in pi.dev within a podman container, using openrouter) start installing various python packages (and uv!) within the environment (not as root) when I asked it to code review some fairly basic Rust .rs files that were generally stand-alone (it did very nicely work out and write some stubs for them to build them and work out how they worked).
It only gave up with the weird Python installing stuff when it discovered one of the Python packages needed Tensorflow.
It seems pretty focused and persistent in continuing its initial approach, and I'm wondering if I need to alter some instructions / initial prompts to rein it in a bit...
[dead]
> while the cost of allowing "how do I synthesize the Spanish flu" is approximately infinite
I've heard this sentiment repeated elsewhere, but why? What makes you think that's the case?Under this rationale, every serious HS textbook has "approximately infinite" risk. That's clearly not so. Why is this any special?
[dead]
You are comparing wet work in a lab to writing code on a computer.
When you screw up an exploit, you fail to execute the exploit. Famously, just like software's near zero marginal cost of distribution, the marginal cost of failure is nearly zero.
You can screw up an infinite number of times on your way to a successful exploit.
If you screw up with lethal agents in a lab? You die.
Here's a non-exhaustive list,
Dora Lush died after accidentally pricking her finger with a needle containing lethal scrub typhus while attempting to develop a vaccine for the disease
A 23-year-old laboratory assistant at the London School of Hygiene and Tropical Medicine, was infected with smallpox after observing the harvesting of live smallpox virus from eggs without isolation cabinets at that time. The assistant was hospitalised and before being isolated, she infected two visitors to a patient in an adjacent bed, both of whom died. They in turn infected a nurse, who survived
Ebola laboratory infection by the accidental stick of contaminated needle in the United Kingdom
Researcher Nikolai Ustinov was lethally infected with the Marburg virus after accidentally pricking himself with a syringe used for inoculation of guinea pigs. The accident occurred at the Scientific-Production Association "Vektor" (today the State Research Center of Virology and Biotechnology "Vektor") in Koltsovo, USSR (today Russia).
"lethally infected with the Marburg virus after accidentally pricking himself"Anything lethal enough to kill other humans is lethal enough to kill you.
And if you don't know what you're doing — and for this argument you're saying this person has to ask a LLM "how do I spanish flu?" then they definitely don't know what they're doing, the number of ways you will die far outnumber the ways you can succeed.
And this, of course, doesn't even cover the cost of equipment, the precursors, sourcing the highly specific materials needed, then setting the equipment up... etc.
The same is true for the Bosch-Haber / Haber-Bosch process, which famously made WW1 possible. Every HS'er learns about the process and the steps. Steps that were classified once upon a time and were the subject of negotiation at the Versailles.
Does that mean a HS'er (or any adult) can set up an experiment that works at 177 times the pressure of the Earth's atmosphere to do anything at any scale without significant infrastructure and help?
The people who can do this are domain experts, and they've been able to do this with COTS stuff since the 1990s, at the very least, for a price of around $2M – https://en.wikipedia.org/wiki/Project_Bacchus . And those people don't need a LLM to tell them what to do. In fact, they're the exact people who'll have access to unrestricted versions of these LLMs.
And from a security perspective, I would bet good money that flooding the FBI's tip line with junk about every teenager trying to learn "what be a mitochondria" does more harm to the effort of finding people who could be planning such a thing than it helps. It takes more resources to go through the mass of false negatives that have now been created as matter of policy.
These experiments have been run. The fictional scenario of someone learning how bioweapons work and conjuring up a plague isn't real and it hurts humanity as a whole to impede the sciences over it.
Because what someone can flail around in / do is learn about immunology / try to "cure cancer" with a LLM and hopefully get started on a long career in medicine. Or, a discovery that matters.
Because in those cases, if and when they do end up at a lab, screwing up doesn't mean death. Just tons of wasted time (and money). And they will fail / screw up. Just look at literally every undergrad in any lab and the expensive messes they create.
-
And last, but not least, yes. Teenage hackers have been a meme for decades.