I've been casually spot-checking other animals in other vehicles, because my absolute dream situation here is to catch an AI lab that's demonstrably better at pelicans on bicycles than other combinations.
Catching a lab cheating specifically on my one dumb benchmark would be really funny.
Dylan's methodology here - generating 1008 SVGs across an 8x6 combination - is significantly more robust than anything I was considering.
His conclusion:
> Nothing jumped out at me. I couldn’t find a case where the pelican-bicycle images looked noticeably better than the rest of that model’s grid.
Because otherwise using a LLM to generate complex svgs is pretty niche and what I thought made this a good benchmark when it was new - generalized programming and spatial knowledge.
Obviously image gen in svg format is not a particularly hard problem if tackled directly on its own.
Instead, it feels like a more appropriate benchmark for the original purpose would be to come up with new, novel problems each time, and compare across all models (including previous ones).
Throwing my hat in the ring: Generate a pelican shaped crossword where all the clues are related to bicycles.
(Haiku 4.5: https://imgur.com/a/N112Nxo, I'm trying some others but it's very slow! Opus has been at it for about 20 minutes.)
"Simon Willison, among other things, is an advocate for the inclusion of pelican geometry in LLM training datasets."
So I tried "Make an image of a Djibouti cab with a camel sitting in the passenger seat. Give the camel no human anatomical features." Still bad.
Two years later the graphics and perspectives of the tools are so much better. But that isn't the big story. What really stands out is the models are no longer adding human female anatomies to camels and octopi. The gates and filtering based on instructions have matured enough so that the pelican/bicycle deductive reasoning puzzle is less problematic. But until an LLM anticipates something like impressionism from Parisian artists rebelling against the rules of the French Academy, continue to reserve a place for human artistry.
tl;dr simonw, your spot-check nailed it just as well as any extensive methodology. LLM's no longer need to cheat this part of the test (better to hack the question than the tool).
And since we have established this silly routine once, I must keep going and ask - yes, you have quite a collection of bona fide pelicans you’ve seen and photographed.
Have you physically seen all the other animals you’ve evaluated as well?
I was born in Soviet Russia, so trust but verify and if still around, maybe there pelican make LLM draw Simon ride bicycle and notice Python code improve.
Similar thing happened when TPC came up with SQL benchmarks.
If you're not good at TPC, your engineering team is no good.
If you're good at TPC, then (as a customer) we will actually include you in a bake-off benchmark for our specific problem.
Winning on it is the price of admittance into the game, especially in a crowded market.
But how narrowly you benchmarket matters, you can't just hard-code that specific scenario & not fix anything adjacent while you're at it.
For example when it comes to GPUs, the "Quack3" (sic) benchmark on ATI cards comes to mind.
They're mostly not very funny though.
Snakes on a plane, weasels on a diesel, spiders on a glider, baboons on a balloon, goats on a boat.
And that's how we end up with PelicanSkynet
And a man without religion is like a fish without a bicycle.
[dead]
> However, facing right is common: 60% of all 1,008 images do it. How common depends on the animal and the vehicle, and bicycles are one of the two vehicles where it’s strongest
Of course the pelican on the bicycle is facing right. The drivetrain on a bicycle is on the right side. If you want any representation of a bicycle that shows the drivetrain you're going to show the right side of it if you want to do so without the frame occluding it. It's an excellent bet that their training data reflects this.
Citation: https://www.rei.com/c/bikes
Edited to add:
As near as I can tell, all of the bicycles are shown facing right, regardless of the direction the animal is facing (GPT 5.6-Terra, Sample 1/3). Also, in every case where the rider has legs (i.e. not the whale) both of the rider's legs are on the right side of the bicycle. This suggests a pretty serious lack of actual understanding of how a bicycle works.
1. https://simonwillison.net/2025/Nov/13/training-for-pelicans-...
There is a bunch of guidance online on how to photograph bikes, and every sales image of a bike will be from the right. You can anecdotally observe this by google imaging 'bicycle for sale'.
Take a look at the GLM 5.2 and Deepseek V4 "animal on a plane" examples. In every case, the animal is standing on top of the plane, a clear misunderstanding of the concept of "animal on a plane"... with the exception of the Otter. The otters are sitting in a seat on the plane, looking out the window.
That's Ethan Mollick's "Otter On A Plane Using WiFi" image benchmark.
https://www.oneusefulthing.org/p/the-recent-history-of-ai-in...
(Sometimes the Racoon is sitting inside the plane as well, but the racoon is a common backup benchmark. I'm surprised it wasn't also holding a sign saying that it loves trash.)
Also, Grok seemed to really really enjoy "whale on a plane" in that second round, and kudos to GPT Terra for deciding after 3 rounds that the user was terrible at spelling and generated "Antelope On A Plain".
EDIT: I promise I'm a human, but I did just notice my "that's not x... that's y" construction at the start. I am rather Claudepilled — my apologies.
Looking for evidence of the same, but with another twist: checking if the models would choose to create a pelican on a bicycle, if no specific bird or method of transportation was specified.
My version of it: https://www.modelbias.ai/pelican-on-a-bicycle-test
https://www.modelbias.ai/pelican-on-a-bicycle-test/result/12... https://www.modelbias.ai/pelican-on-a-bicycle-test/result/12...
Exactly--and you have to ask yourself at this point what "maxxing" really means, since "get better at drawing SVGs" is a useful skill.
Pelicans aside, we need to remember that benchmarks are the only good quantitative way we have of comparing models. If someone has complaints about “benchmaxxing”, please ask them to contribute a better benchmark! It is valuable work and very appreciated.
Me: how many P's are in the following text? [Pasted text]
Claude: There are 14 P's, all lowercase (no capital P's)
Me: how many in "strawberry"?
Claude: there are 3 R's in the word "strawberry".
Me:how many P's are in the following text? [Pasted text with 10 P's]
Gemini: There are 9 "P"s (1 uppercase P and 8 lowercase ps) in the provided text. [List of words except the one missed]
Me: How many in strawberry?
Gemini: Something went wrong (1096)
> Direction: All 21 pelican-bicycle images, across all seven labs, face right. No other animal/vehicle combination does that.
> However, facing right is common: 60% of all 1,008 images do it…
and the tables in “Evidence #5” to be anything but evidence the models have likely trained on pelican on bicycle data more than others.
The data clearly shows:
- 100% pelican on bicycle facing right
- significant skew to the right for bicycle-like vehicles
- significant preference for right facing for birds
Averaging those extreme results to “60%” to make it sound like it’s pretty fair because it’s close to “50%” isn’t statistically sound.
The methodology is generally unsound. There is no actual scoring with a well defined rubric, it’s just vibed with a single model (GPT 5.6 Luna).
The “not better at drawing” evidence are equally hard to take seriously when there is no clear, non-subjective indication of what better or worse is.
https://x.com/CrimeDecoder/status/2080008114615537766
Can see the images for ChatGPT/Claude (Sonnet 5), and Gemini are all quite bad.
Jagged edge of LLMs. How do you explain being able to generate very complicated shapes in the Pelican example but cannot make a much simpler icon without just alluding to it is in the training data?
However thinking about it...if someone asked me to manually create an SVG, or hell even draw a quick doodle on a bit of paper of the same, I'd still probably be much slower than an LLM and potentially end up sketching less accurate anatomy than the machine.
I think the general "organic task" stuff has been mostly sorted out, but in personal and professional experiences using AI to try to _do_ something, I've found less so recently problems with hallucinations and moreso problems with attention.
For example GPT5.6 still has issues where if I provide it with a list of documents and then ask it to raise questions from that information. Then provide it with additional documents that answer some of those questions and ask it to summarise which outstanding questions there are again, it still asks questions that have become irrelevant with the additional documents - but when this is pointed out it knows exactly what to do and produces the correct list of outstanding questions.
I'm sure frontier models are doing all sorts of crazy stuff with attention already, but it seems to me like we almost need some hierarchical attention mechanism like KVL (with Level added) so that it's aware not only of semantic connections between tokens in the context but also of where there are gaps, missing links to assist the model in becoming aware of its own attention span (I guess).
Image generation models nowadays can easily generate a photorealistic pelican riding a bicycle, where the bicycle has perfect structure. But it is, of course, only a raster image.
It seems that we're missing a kind of step to decompose an image into a list of instructions (say, SVG paths, or even brush strokes with a real brush) to reproduce it properly. Doing so would probably need a true understanding of the structure of the scene, which is something that AI still struggles with to this day.
Same pelican grid + MacBook Pro in 3D. Also with end cost.
https://s3.eu-west-1.amazonaws.com/images.dylancastillo.co/p...
https://s3.eu-west-1.amazonaws.com/images.dylancastillo.co/p...
aren't some LLM going to digest that thread at some point and indirectly learn from it?
basically my point is originally this was a good benchmark because it was an absurd never-seen-before thing, but now that it is in content, some models are going to get a benefit in education?
you'd need the "AI" equivalent of an old-school "google whack", something with no previous results
So they aren't pelicanmaxxing but they are benchmaxxing in a way. The benefit of the pelican was originally that uplift on the pelican signaled an overall uplift on intelligence. I don't believe that is the case anymore and it is just another jagged edge of model intelligence.
On the flip side, GPT 5.6 Sol did a pretty convincing render of a burglar eating salami.
I'd be curious to see how the other models on random things that are completely tangential to pelicans or bicycles.
Obviously they're all a bit cartoon-y, what else do you expect from SVGs. But I'm not convinced you could find a single human on Earth over the age of 4 who would seriously give the vehicle in GPT/1/whale/plane a 5/5.
Browse through the options a bit and the rest is not that much better. Grok/2/cat/plane, one of the more accurate planes, got a 2/5. For the most part, vehicles entirely missing do get a 1/5, except for whatever it is in Gemini/1/heron/plane scoring 4. Animals inside planes get completely random vehicle scores I guess.
The cats are all orange, except for a few of the skateboard cats that are black. I'm sure there's nothing to read into there...
Well, I've convinced myself that the next effective test of multimodal models will be whether their judgments of LLM-generated SVG airplanes are anywhere close to reasonable.
I'm curious if some of the animals / vehicules might force the models to use more tokens than others, and I could not find the token counts in the shared data, is it possible to publish it please? :)
Since the tests can be generally gamed with directing descriptions at it non-deterministically, there's a greater chance the questions solution can be found.
Of course, hopefully the models are instead adding patterns and types of questions as well and it makes the models more capable, but it may be limited in how it transfers to other types of questions in breadth or depth.
However, I'm very surprised that most of the models make the same sideways mistake with only some of the animals, and they do it consistently.
With most of the models, cat, raccoons, and otters are almost always riding sideways. Why is that?
In the last century, as spec tests for C and C++ compilers, databases, Java application servers became trendy, all vendors were optimising for great articles on the respective technical magazines.
It made me realize that for this time around, we’re not at the center anymore. The future of LLMs is the stuff of nations and AIG is what the labs actually care about. They aren’t pelicanmaxxing just as much as they really aren’t revenue/margin/marketing maxing. They just want our attention and ideas so they can show growth and acquire FLOPS.
The apparent fact that they aren’t catering to this community (who frankly decides what goes and stays in prod) leaves me feeling a bit defeated somehow. And also awesome?! Like there is a culture here that runs deeper than any technology and it has fought hard to maintain its identity. Props to dang and all for keeping the astroturfing so imperceptible that I can say this.
In any case, a great little piece of citizen-science dcastm. Lmk if there is a way I can chip in towards token costs.
Even if using LLM is the only way to do things at scale it does not mean that it's always the right tool.
Your next iteration will need different animals and different transportation options. You'll run out after a few iterations.
I think we should stop using pelican benchmark.
> Every single one is a pelican, on a bicycle, first try. When every student gets an A, the exam has stopped grading.
Numerous pelicans and their bikes are clearly horribly malformed. In fact none of the bike frames are correct. Fable and Opus come close, but the top of the diamond is disconnected in Fable's case and the head tube is misaligned with the front fork in Opus's case.
And of course, as the parent post shows, labs don't actually seem to be training on the pelican bike case.
[deleted]
very nice approach to test it and might be a nice way to "grid search" evals in other use cases perhaps.
He was asking the question - do we see gains across other tasks? The underlying question was: Is the additional attention given to this specific task creating a false impression of progress?
"Any observed statistical regularity will tend to collapse once pressure is placed upon it for control purposes."
Or the more pop layman version
"When a measure becomes a metric/KPI, it ceases to be a good measure."
Story time, I live in Argentina, and we don't have Big Macs, the main Mc Donald's brand, here, because during the CFK presidency, one of her tactics was to Goodhart economic metrics. Even the informal obscure ones like the [Big Mac Index](https://en.wikipedia.org/wiki/Big_Mac_Index), I don't know the precise details, but the Big Mac ended up being a very cheap item, like 2 or 3 times cheaper than actual menu items, but it was never on the advertised menu, and it also ended up being very small compared to the other burgers, so it wasn't even like a hack, a shrinkflation type of deal.
But hey, anyone who read the Big Mac Index table would never find Argentina at the bottom of that list along with a couple of other countries with bad brands, so the ploy worked. And now we live with the aftershock, the brand never really turned around, other brands with ridiculous names took over it like the McTasty, which makes me sound like that skit from Tarantino's Pulp Fiction.
The right answer here is to ask a LLM to create a scene similar in quality to those, but completely out of distribution.
I asked GPT 5.6 Sol to give me a pelican playing football on San Siro while smoking a cigarette, in AC Milan's t-shirt. While this sounds like higher complexity of a problem, the generations from current models often include additional details like scene composition, scarf, etc., I don't ask for, so I wanted to see what here is memorization vs. composition skill.
"write svg code of a fish playing football on san siro in ac milan's t shirt, with raybans on and a cigarette."
Try that on GPT 5.6 Sol, Fable, or whatever other model. It's chaos.
Whatever. Doesn’t really matter much. My next thought is, I get that requiring an SVG is adding an extra layer of complexity as far as the art goes, but why is nobody talking about that the actual art is absolute trash?
I get it. It’s basically a meme at this point and it’s a fun game to play with the models. But my thought is it should be illuminating to anyone who is an artist that LLMs are still a long way off from taking your job :)
I don’t have any reason to believe they are gaming the benchmark, just saying. I do find the idea of a data labeller having to generate thousands of svgs of different animals on different modes of transportation quite funny though.
Tomato, tomato
> But again, some combinations might be just harder to draw than others.
> To account for that, I fit a fixed-effects regression on all 1,008 images: score ~ lab + animal × vehicle, plus per-lab interaction terms for pelican, bicycle, and the pelican-bicycle cell, with robust standard errors. The animal × vehicle terms absorb the inherent difficulty of all 48 combinations. The interactions measure each lab’s benchmark-specific boost relative to the average lab, with confidence intervals.
[dead]
[dead]
[dead]
[dead]
[dead]
[dead]
Test the LLLM against things you want it to do.
Asking questions that are absurd is like interviewing developers and asking absurd questions on the grounds that it tests creative and critical thinking.
Remember these Microsoft interview questions designed to identify the best developers?
"If you could eliminate one U.S. state, which one would it be?"
"How would you move Mount Fuji?"
Absurd interview questions have an air of legitimacy due to the quasi sophisticated justifications put forward for why they are good tests.
Absurd interview questions are not good tests of people or LLMs.
Relevant questions are good tests.
This makes me think AI companies are using chat history to train the next model.
It may be that AI performs better on such code.
(don’t tell my boss.)
reminds me of this Key and Peele skit
If you stick to the benchpress, it's just "benchmaxxing".
Minimizing peaks is probably not a good strategy generally. Even low peaks may have some benefit ("if you know your problem is in this domain and performance is critical, this tool offers a 3% advantage").
Minimizing troughs may be more attractive. People generally seem to react more strongly to negatives and competitors can devise benchmarks which emphasize one's troughs.
While a universal expert would be convenient and broad knowledge aids some forms of creativity, specialization has substantial advantages.
Trading max performance (primary metric of concern) for some improvement in a secondary metric often makes sense and that could be an interpretation of "minmaxxing", reducing the over-emphasis of a single metric which would otherwise be maximized.
AI image generation suffers from this more generally. You can generate pictures of pelicans, sure. Newer models clearly generate images with more pelican-ness than before. But all of it is still uglier than sin. Drawing things accurately is one thing, making results that someone might actually want to use (without embarrassing themselves) is something else.
As for conventional diffusion-model stuff, I happen to think there are some pieces of AI art that still look really good even knowing they're AI.
There seems to be a very strong correlation between models that are good at SVG and models that are good at 3D CAD.
Anecdote I know, but there does seem to be generalization going on here.
I haven't really tested Opus 4.8, but 4.7 wasn't nearly as good as ChatGPT 5.5.
...which.. hmm I dunno if they are same or not
That said, I think this would correlate relatively little with general programming ability. They're not unrelated, of course, but being able to generate code that paints an accurate + esthetically pleasing image is quite different from generating code that achieves a non-spatial goal.
Why do LLMs need to be able to do this as well, but worse, slower and more expensive?
I'm almost about to post that xkcd 810
I don't understand the truth to this point, and there's a clear difficulty in establishing structure, a thing that you don't need to actually do with raster image gen, right?
Like if you say "please give me 9 circles" you expect circles in the image and not just a bunch of pixel that are vaguely circle-like, right?
That simonw is causing labs to do extra fine-tuning runs for this seems highly probable :)
SVG has its neat SVGO tool that optimises SVG code, this is very clever but it won't spot easy wins such as when a group of elements can be mirrored, or an element re-used with a transform. Neither will it use the full smorgasbord of features. SVGO is not AI, however, it does a fab job of taking bloated files from Adobe Illustrator and getting something good to work with.
The bar for SVG is really low, mostly just paths, compressed and not human readable. SVG should be a human readable format, so 'circle radius 10' rather than two hundred points at six decimal places to draw the same circle.
You have to RTFM to do cool things with SVG and there is a lack of appreciation of the format amongst developers and designers.
Mission fucking accomplished. https://xkcd.com/810/
We see the same problem in database benchmarking.
The good thing about the deliberately-non-real world pelican case is that it gives a general impression of how much the model is improving because it's not likely that it's being specifically targeted at it, rather than a 'real world case' which might have been specifically optimised for.
https://www.gianlucagimini.it/portfolio-item/velocipedia/
On the other hand, I notice that the prompt says to draw a pelican riding a bicycle, implying motion... and since most of us read left to right, I think it's usually natural to draw an object in motion moving left to right as well, which means the bicycle should be facing right. So maybe that specific setup is more natural here.
Either way, for humans bicycles are actually really hard to draw from memory. In fact, I substitute teach, and sometimes as an activity I have my students draw bicycles from memory in 60 seconds. Most make pretty serious errors, usually the frame or chain connections: they can tell it's wrong but still can't draw a more correct one. I use it as an object lesson about the difference between recognition and recall - most students never realize that much of their studying can end up being the former, when tests and life almost always ask for the latter. This helps explain why many students go from "that makes perfect sense" when going over review problems to a total mind blank only a few minutes later (especially in math!).
Independent of what humans may do with the same prompt, training datasets will be full of traditional bike photos, which will all be drive side.
It’s also one of the easiest yellow flags to look for in used bike listings (not low-end, but anything enthusiast level). If a bike is photographed on the wrong side, there’s a non-negligible chance it’s stolen, because the seller is presumed to know better if they’re an enthusiast themself.
Looking online, the kickstand is apparently on the left to avoid the gears. Based on the other comment about bike photography, it's interesting the same design choice makes humans and cameras/LLMs see bikes from different sides.
an... understanding? there is no understanding here at all
https://www.cyclingweekly.com/news/japan-unveils-new-olympic...
(I'm not saying that they did that. I'm just saying they can.)
> Using a single LLM judge for scoring. Every score here comes from one model, GPT-5.6 Luna, looking at one image at a time. I didn’t do much alignment and didn’t check how often it agrees with itself on a re-run.
Having used a similar setup (with previous gen LLMs) to evaluate the 3D models that my product[0] generates, it turned out there was no correlation at all. LLM judgments were very much random and I assume judging SVGs is not that far from judging 3D models. I guess I have to re-test this with current gen.
This post proves that hasn't happened yet, either. Although maybe the bad results posted online are being trained on and that explains the UNDER performance.
The very best svg pelican on a bile generation model. Just for laughs.
When LLMs use the pattern, they are often setting up a straw man and then knocking it over.
In my tests it did create bicycles the most, but this is just a general bias I believe, as tested here: https://www.modelbias.ai/prompt/transport
If you train towards the test, you aren't necessarily improving overall fitness, but you are destroying the value of that test over time because you're decreasing its correlation with overall fitness.
I don't think I'd go that far!
When someone says a model has been benchmaxxed, what they really mean is that it performs better in benchmarks compared to their real world experience. That's a real thing, I've certainly experienced it with some models.
...my take is that some things in life just resist quantitative measurements. Who is the best job candidate? What is the best programming language? Add AI models to the pile.
[dead]
> Here is what Qwen3.6-35B-A3B via Openrouter provided for a sloth riding a skateboard: https://imgur.com/a/Dy8fvR5
> Like Grok 4 Fasts attempt at a mushroom in a rowboat, it is barely recognisable as anything despite both Qwen3.6-35B-A3B and Grok 4 Fast having no issue with more popular (i.e. benchmarked) examples. [...]
> And here is Opus 4.7 [which simonw claimed to provide a worse pelican vs Qwen], again via Openrouter: https://imgur.com/a/Qus1Enf
Anyone who hasn't witnessed such deltas either hasn't looked at enough examples, a sufficient variety of models, or both. And they are, unfortunately, not limited to "SVGMaxxing", but a wide range of evals.
Don't judge the dog's technique. The miracle is that it's dancing.
I would guess most programmers struggle to create SVG icons - I don't find it easy. The average person even more so.Are we best to assume an LLM is a blind programmer? Any HN comments from blind programmers tasked with creating SVG icons? Only relevant comment I could find from ctoth was about accessibility: https://news.ycombinator.com/item?id=7185771
Projecting how you think onto what the LLM is doing or should be doing, is probably a mistake on your part.
I recently spent a little time trying to understand exactly why Gemini was misexplaining $X. $X = {why the generated LaTeX visually didn't match what it was asked to do}. It was enlightening.
It just might feel foreign to human who does not have a SVG trained head-space.
As such, modern LLMs still kind of suck at generating pelican bike SVGs with obvious errors:
* some omitted the bottom of the diamond which connects from the pedals to the rear wheel
* some added an extra connection from the pedals to the front wheel, making it impossible to steer
* none could align the head tube with the fork
* none added a correct offset to the fork
* none could generate the chain properly in a way that attaches to the two sprockets correctly
... whereas these errors do not appear in the raster image. (To be fair, the raster image has other weirdnesses, like the bird having arms and only one leg)
If we could harness the sort of internal thinking that must have happened when it generated the highly consistent raster image, but make it output SVG instead, we'd get much better SVGs. So this would be indeed "visualize in their mind how to draw a picture" before outputting the SVG.
[1] https://chatgpt.com/s/m_6a611c29c02481918fbb0f165eb83594
It is the reinforcement learning that produces more tangible results with less data, but it is something that the AI labs specifically selects and it is not picked up unknowingly
The problem isn't the test, its that is a public test.
Simon has previously said he has a list of secret prompts (at least one of which he "burned" as a demonstration a while ago). That's what makes it a good test - his commentary on the public test is something of a proxy for non-public tests. This makes it a good benchmark.
I'll look for a better image host in the future. I guess the economic incentives makes them all turn bad after a while.
You could train a model to do something a bit different like a pencil sketch, or vector graphic sketch, of something given a photo of it, and expect that to be a generalized skill, but if you are asking the model to do it "from memory" then memorizing a pelican is no substitute for not having memorized a horse.
To be clear McDonald's didn't pull out of the country, like in many other countries, it adapted to the local market. Same way as in india it serves non meat variants owing to the high vegetarian population, consider that mcdonald's operates in Venezuela and China, and operated in Russia up until the Russia-Ukraine war. It takes a lot for MCD to pull out of a country.
As for the exact mechanisms of presidential price control on MCD Argentina, I don't have the specifics here, but I can get pretty close. There have been 2 broad mechanisms to exhert price control during CFK's 8 years of presidency and during his Husband's 4, let's call them official and unofficial.
The official mechanisms would be passing laws or presidential decrees (DNU) that don't go through congress, as well as influencing regulation of executive ministries/departments like central bank norms, exchange rates. Some measures like 'precios cuidados' were placed for this very specific purpose, it's very possible Big Macs were under this specific scheme, I do not recall, I was a bit young.
The unofficial methods would be less public, but well known, in the food industry it was especially common, it's well established that food are one of the first and most common targets of price controls. I have heard direct accounts from family members about high ranking government officials setting up meetings with producers in food markets to give orders of lowering prices, one going as far as brandishing a firearm by placing it on a table while discussing the subject.
Writing this out loud I realize that this explains the later over-correction of argentina that allowed freemarket capitalist libertarianism to rise. The optimal strategy in democracy seems to be polarization, so both extremisms seem to symbiotically feed off each other. Not a country of moderateness this one.
Yes; what's wrong with that?
Do you suppose that it doesn't test those qualities?
What sufficiently hard, but useful, problem would you ask the model for?
I agree, it is ridiculous to ask an LLM to replace an artist.
Yes, deliberately so.
It was never intended as a meaningful benchmark. The surprising thing was that for the first ~12 months performance on the stupid pelican benchmark did seem to correspond to the performance of the models on other tasks.
That pattern no longer holds - Fable 5 and GPT-5.6 have both been out-pelicaned by lesser models now.
It could well be that the exact domain matters more than the bigger picture concepts, like 3d. One of the ever fewer reminders that this tech is still just fundamentally a token prediction algorithm.
Mostly I have been using GPT 5.5 and now 5.6
That is notable because I do almost exclusively use Claude for coding.
LLMs do a great job because they understand both code and SVGs well.
Edit:
An example for a synth I'm buulding: https://imgur.com/a/U694Ek7
The irony of having to post it as a PNG isn't lost on me...
Why on earth would you let an LLM do this when it would take you 10 min to do this in Figma or Inkscape or even just Word
I haven't read the code.
Why would I do it in a slower, more difficult way for something that's going to be outdated in 2 hours?
[deleted]
This design is trivial. I admit it would be hard to achieve in Word (or at least for me because I don’t know how to make a diagram in Word any more) but Figma and Inkascape are made to do these things, and have optimized UI for that (personally I would have just used mermaid though).
I think you may have lost your faith in human capabilities just a little bit if you think drawing stuff like this takes any time or effort at all. Compared to designing and architecturing the system, drawing the diagram is trivial. Now I know that my parent did neither but it seems like they vibe-coded the whole thing. I’m sure they will end up with a fun little toy from the whole endeavor they can play with for 2 weeks before abandoning. Maybe the author will even feel bad about the carbon footprint of this whole exercise and buy some carbon offsets to make up for it.
That's exactly the idea - except it's more like 2 hours before I prototype the next version.
> Maybe the author will even feel bad about the carbon footprint of this whole exercise and buy some carbon offsets to make up for it.
The passive aggressiveness of this is perfectly weighted and admirably phrased.
Where I'm from data centers help the renewable mix by subsidizing transmission from other geographic zones. I'm actually improving the environment by using it.
[1]https://ourworldindata.org/how-much-energy-do-data-centers-a...