Going on baseless speculation, the lack of accompanying pro models with these flash releases either means: 1) the model is too big to be economical, 2) google doesn't have the compute to serve the big model, 3) their big model has too many alignment issues to serve to the public.
edit: looks like benchmarks are up on https://artificialanalysis.ai/models/gemini-3-6-flash. It's solidly middle-of-pack. However, if you want to be most fair to flash, look at the intelligence vs time per task and intelligence vs outputspeed benchmarks. This is a very fast model.
edit 2: I use antigravity from time to time and in my experience, 3.5 flash is an underrated model, so long as you know what it's good for. It's very good at frontend (much better than gpt 5.5) and it's fast, so it's a great tool for iteration. I expect 3.6 to be no different.
For small models (which are probably distilled from their big ones) you can serve them economically all the time and not hemorrhage money.
For smaller models, you're competing with DeepSeek V4 Flash. (Which I think is a 284B A13B?) Subjectively, this feels about as smart as Sonnet 4.5, give or take. And it costs $0.09/$0.18 on Open Router, compared to $1/$5 for the latest Claude Haiku. See https://openrouter.ai/deepseek/deepseek-v4-flash#providers The developer antirez of Redis fame uses this as a local coding model.
DeepSeek did some extremely clever research on hybrid attention to get the prices that low, reducing per-user context cache sizes dramatically.
So, no, when it comes to low-price models, the US models probably can't sustain their current margins there, either.
Fast, light weight, ok intelligence. Perfect for serving 20B+ prompts per day mostly surrounding banal human things.
OAI and Anthropic's cloud spend can cover the revenue gap, as Google is already capturing a large chunk of those guy's revenue.
4) googles big model just performs worse than K3 and GLM so they choose not to embarass themself.
Like I love Gemini and use it a lot to one-shot whole MR with huge contexts, but its just much worse when its come to tool use and agentic coding.
they have search,youtube,android,office suite like gmail,maps,spreadsheet etc
coding is the least of their problem/priority
From the outside they look like they're behind in terms of frontier models, but I think they might be the best positioned to not go out of business when the bubble pops.
Also look at the fact that they've been able to deploy AI-assisted search at google scale. It must be another order of magnitude larger (at least) than the model deployments for OpenAI and Anthropic.
Of course unless you're inside Google it's impossible to know for sure.
I was already impressed by how fast 3.5 Flash was. But I've never compared it to other models in its class for coding.
Why? Coz models in that class are not very useful to me. Time saved waiting for responses usually just turns into time wasted replying to low quality responses.
Google need to release a Pro model ASAP. I am skeptical of the "maybe they don't have the compute to run it" thing. Anthropic were (probably) in that situation with Mythos and they announced it anyway - that's the obvious play for investor relations as well as hype for your product.
It used to be that you could find the edges of the training set pretty easily. No longer.
He says there are many similar vendors and teams with thousands of people in India and other countries.
This implies that normally models are aligned and there is merely a number of issues to fix.
Rumors say 4) it didn't perform well, especially in coding so has been delayed
They are likely deliberately avoiding the SoTA race for a few reasons:
1. Their best models are marginally better than current SoTA releases. 2. They'd like to let Ant/OAI make mistakes with safeguards / let them get the regulatory heat. The unknown unknowns are huge with SoTA models (eg OAI accidentally hacking huggingface) and they are protecting their reputation. 3. They want to encourage companies to become cost conscious because they can likely win on price in the long run. Getting market share in "quantity beats quality" workflows forces companies to establish processes to choose the "cheapest acceptable model", which is a good environment for Google.
It seemed for a time that Google had finally gotten the ball rolling, but I'm doubting that more and more as time passes. We'll see what happens with 3.5 pro I suppose.
They literally forced me and my company out of Antigravity by phasing out AI Ultra subscription without any proper product follow-up. Antigravity IDE cannot even have poweruser subscriptions now from Google Workspace an Gemini Enterprise Agent Platform cannot be attached to Antigravity IDE.
Gemini Enterprise Agent Platform has an incredibly abysmal setup process, and if I want to limit spending per-user I have to create projects per user. The fact that you cannot activate Anthropic models on it if the billing still has free credits is almost a joke.
I was a big proponent of Google and Gemini, but they left us reeling with their abrupt product decisions. Forced us to buy $200 subscriptions directly from Anthropic/OpenAI.
https://tools.simonwillison.net/markdown-svg-renderer#url=ht...
2.5 Flash: $0.3 / $2.5
3.0 Flash: $0.5 / $3
3.5 Flash: $1.5 / $9
3.6 Flash: $1.5 / $7.5
---
2.5 Flash-Lite: $0.1 / $0.4
3.1 Flash-Lite: $0.25 / $1.5
3.5 Flash-Lite: $0.3 / $2.5
As is, they are thoroughly outclassed for most usecases. I will say the one area where i do see Gemini punching above its weight class is in tasks that are effectively "Google this for me" / knowledge stuff. So it does have a role, and I do use it. So while I think Google is still in a strong position overall, they are really stuck as a tier 2 AI player right now with text models. They are tier 1 in bio, images, and video.
> Beyond today’s releases, Gemini 3.5 Pro is currently testing with partners and we plan to make it broadly available as soon as it’s ready.
> We have started our most ambitious pre-training run yet, for Gemini 4, and are excited by the progress.
Hopefully 3.5 Pro is soon, and that Gemini 4 can be here end of year and finally have an updated knowledge cutoff.
I have a very price sensitive workload that used to run on flash 2.5 lite - it's deprecated now.
The replacement 3.1 flash lite is a lot more expensive, but now also has a sunset date.
3.5 flash lite is even more expensive.
So the price is rising and you have no choice but to keep paying more and more.
Anyone have any good alternatives?
I tested Jules and while the idea is good in theory, I found the model's intelligence to be very lackluster.
I wonder if there is something with their TPU cycles that makes them want to postpone training a new model. My guess is that they have been on the same base model for 6 months and they may have waited for the next gen TPUs to train Gemini 4, which greatly limits how much intelligence they can increase and forces them to do cost efficiency increases.
That being said, it seems that Gemini is still the best image analysis model, so hopefully 3.6 flash builds on this even more.
1. Their AI efforts are very fundamental research oriented. They are really good at it.
2. Their productization sucks. The end products gets little attention compared to competition. It can be canceled at any time. You should never build anything around Google only APIs, AI or not.
Has anybody found any models better at image or audio analysis?
gemini-2.5-flash-lite: $0.10 input / $0.40 output
gemini-3.1-flash-lite: $0.25 input / $1.50 output
gemini-3.5-flash-lite: $0.30 input / $2.50 output (a 6.25x increase over 2.5!)
Now watch them deprecate Gemini 2.5 Flash-Lite in the coming months...
This plus the Vertex, AI Studio, Gemini, Antigravity. It's honestly too confusing to use. I need to use Gemini just to decide on which platform and which model to consider.
0: Famous Kurian Tweet: "We're announcing Duet AI for Google Workspace will now be Gemini for Google Workspace. Consumers and organizations of all sizes can access Gemini across the Workspace apps they know and love. We're introducing a new offering called Gemini Business, which lets organizations use generative AI in Workspace at a lower price point than Gemini Enterprise, which replaces Duet AI for Workspace Enterprise."
https://artificialanalysis.ai/models/gemini-3-6-flash?intell...
Rough patch for google ai
Fable 5 still wins on detail with no visible errors, but it's close. And this isn't a memorized pelican;
https://playcode.io/blog/macbook-svg-benchmark#gemini-3-6-fl...
It is also cheaper than 3.5:
> This enhanced efficiency is also combined with a lower price than 3.5 Flash. At $1.50/1M input tokens and $7.50/1M output tokens, 3.6 Flash reduces the overall cost per agentic task, making agents more cost-effective to build and run.
Otherwise, this news feels like a tiny incremental improvement on Gemini Flash series to make it more efficient with token usage, subagent and cost. Nothing big.
Regarding their benchmark scores on CyberGym, I wonder why they didn't compare their 3.5 Flash Cyber model with Fable 5. I mean they included Mythos and GPT-Cyber, so why not Fable 5 too?
They also mentioned Gemini 3.5 Pro is in testing and its about to become available very soon. Another thing maybe worth discussing is the announcement of pre-training Gemini 4. Sadly, not much technical details to discuss on. Many comments in here seem to mostly be about how Google is behind the others, but honestly, is it really worth the investment to be #1 in Artifical Analysis every week?
Ever frontier lab lived it at least once : missing the frontier by a few months triggers extremly negative reactions, then you take back the lead for 2 weeks, and the hype cycle repeats.
It's quite a fun game. I click into the comments and see if the roulette wheel was right.
It's a one thing to research and improve the model, but if they ignore the ease of access and multi-availability of their models in different ways they are going to fall behind again.
That said, the speed looks really good. I think it's competitive with Fireworks's GLM 5.2 Fast, although Fireworks is still cheaper.
If it wasn't for Gemini/Antigravity I'd have to go with a Max Claude plan, as it stands now I can get by with just a Claude Pro plan to get Opus when I need it, whilst using Antigravity as my day-to-day workhorse.
Unfortunately Gemini Flash became too expensive to use as a general purpose model (i.e. for AI features in Apps), luckily there are plenty of cheaper Chinese models to fill that gap now.
It's time for them to start focusing on open-weight models and efficiency. Otherwise there's just a layer of marketing hype and "will it do this?" that has to be cut through for evaluation of each and every release cycle.
Models are getting easier and easier to create. The money, if there's any here, is in the harness the user interfaces with, and the data centers running them.
Spawn 10 on the same problem and have them debate to reach a consensus, you’ll get Fable-like results but 100x faster.
> We have started our most ambitious pre-training run yet, for Gemini 4, and are excited by the progress.
Seems like this is mostly a cost play by Google, hoping this doesn't bring 3.5 Flash capabilities to an end of life, and that 3.6 catches up or gets better.
Plus they are probably running these things on every Google search so saving tokens is a huge win for them.
3.5 Flash Lite is only a hair cheaper than 3.0 Flash, but I think 3.0 Flash is a massively more capable model?
Screw your government! US and Israeli governments should get the least access, but of course we all know they'll be the (only) ones to get full unfiltered access.
"3.6 Flash delivers higher precision with fewer unwanted code edits and reduced execution loops, as seen in DeepSWE (49% vs. 37%)"
So which one is it? 65% or 49%?
Deepseek Pro: 0.435/m 0.87/m
That's wildly ambitious pricing by Google. You can maybe get away with spicy pricing at the SOTA edge but at the lower tiers everything is a lot more price sensitive.
Google was late to coding agents and as-per-usual fucked it up with their crazy project management culture.
Usually Google gets away with it due to inertia, however this time they are paying a heavy price because they missed out on the training data that Anthropic and OpenAI have gathered with claude and codex.
My results [0] put Gemini 3.6 Flash at the top.
3.6 Flash high has same $1.5 input price as 3.5 Flash, but output is cheaper from $9.0 to $7.5.
Google said 3.6 Flash is more token efficient, but in my tests it's actually LESS token efficient[1] than 3.5 Flash, so despite the output price reduction, it still costs more.
[0]: https://aibenchy.com/compare/google-gemini-3-6-flash-medium/...
[1]: https://aibenchy.com/compare/google-gemini-3-6-flash-high/go...
GLM 5.2 is better, also cheaper, and almost as fast.
So essentially, a big L for Google. Combine this with them not being able to produce a frontier model this generation... hmm implications
I'd be low-keying the release if anything, given how lame they are compared to their competition.
What am I missing?
I use 3.1 Flash Lite regularly to classify listings on eCommerce websites. It's great for this task - fast, cheap and accurate.
In fact, it was the single best model we tried in terms of the speed vs accuracy vs price tradeoffs - including the Chinese models.
Of course, 3.5 Flash was more accurate but the 5x cost increase couldn't be justified.
3.5 Flash Lite sounds like it could be a strict upgrade for our use case, without a significant increase in costs or drop in speed.
It's not GPT-6 but it's not trying to be. It's a completely different tool and great at what it does.
Meaning, its predictable with tool calls, wont spin off a million tools/do weird behavior, its reasonable. Even sonnet in a real world decision making scenario is not reliable, or will reason so long its incredibly expensive.
The benchmarks arent catching all the value, and most people have never actually ran an ai agent in a real context that matters
[deleted]
Front page is tedious these days.
1- no comparison with gemini 3.1 pro
2- no comparison with any other model
https://storage.googleapis.com/deepmind-media/Model-Cards/Ge...
we are stealing plutocracy from the jaws of emancipation.
i don't want to live in a world where abundance is guarded and shared among politicians and cronies, whilst the rest are left to rot.
Gemini 3.6 Flash https://news.ycombinator.com/item?id=48993130
For me, Gemini models are the most usable. Claude Opus and Mistral always try to turn queries into one-shot enormous commits, which just burns tokens, time and annoys me for something which is still wrong more often than not.
Gemini seems far better at listening to instructions and giving me what I actually want, on top of using far fewer tokens and wasting my time. Fable is the only model that's come close to Gemini Pro for me.
And as this is about Flash, it's exciting, I find Flash can usually get the right answer pretty quickly and without too much nonsense.
Google watches over the last few months a flat out assault on the Pareto curve from American and Chinese companies. Release after release pushing the boundaries of frontier intelligence and price/performance.
And the response from arguably the biggest AI research labs in the world by headcount is Flash 3.6.
What do you do when you are given essentially unlimited resources and still find yourself falling behind?
...and last time I looked the limits were more generous for Gemma 4 there, but they have been tightened a bit. That's how it goes, always changing.
It doesn't matter how good the model is if you're (mostly) forced to use it in Antigravity - which turns any model into crap.
Wake me up when Antigravity doesn't suck.
> For autonomous subagents with tool calls, code execution, or multi-step reasoning: set thinking_level to "medium" or "high" to prevent premature tool termination.
I just happened to see that in the docs: https://ai.google.dev/gemini-api/docs/latest-model
[dead]
[dead]
[dead]
[dead]
[dead]
[dead]
[dead]
[dead]
[dead]
[dead]
[0] - https://www.ft.com/content/6049a031-9e9b-464c-97bb-414da04d5...
you can check by asking "list notable world events in 2025, only list unplanned" on aistudio. or you can ask for Charlie Kirk, it also does not know. I tried it multiple time to ensure that I didn't not get routed to older models!
> but google has search
irrelevant, without deeper knowledge about cutting edge technologies or latest libraries, all of it suggestions are crap. even you ask it to search it will still use outdated keyword thus only getting outdated information.
in other word, what a disaster!
[dead]
[dead]
I don't think so. US models are very expensive, and not available in every country. I am not willing to pay $50/1M tokens for writing my pet projects.
the reality is one way or another that as long as there exists an alternative that a USA company could serve with the same compute rented from hyperscalers, this represents a threat, even if the extent to which is unknown
https://openrouter.ai/rankings#top-models
Of course openrouter is not representative because most users directly go to the model provider but it still proves your claim is very far from "absolutely true".
I think the reason OpenAI and Anthropic stay ahead in revenues right now is because the models are improving too quickly to reliably compete with them on cost.
But, once model performance reaches a plateau -- they have to at some point, though perhaps years away -- that's when ability to operate compute infrastructure at scale becomes the secret sauce.
The big AI labs are likely safe until models stop improving fast enough to protect them from competition on cost.
This similar pattern has repeated in most technical booms prior to this.
When hard drive technology was improving fast enough that old hard drives were quickly obsolete, IBM could maintain good margins making hard drives. But once hard drives got good enough and advances were slow enough that innovation was not the only factor considered by drive purchasers, commodity hard drives started to take over and IBM had to exit those businesses.
The same is likely to happen once model improvement slows.
On the very link on the top comment of this thread, which I repost here:
https://artificialanalysis.ai/models/gemini-3-6-flash
Kimi K3 is ahead of Fable 5 on several benchmarks.
So basically the angle went from "China cannot ever compete" to "China is six months behind" to "China is six weeks behind" to "China is six days behind but that's because they're distilling" and now you're saying "Yup sure, Kimi K3 is ahead on several benchmarks but you cannot host it yourself so this thing will go absolutely nowhere".
I mean: is it not a bit early to draw conclusions? It's been days since a chinese model is ahead of the very best / frontier US model on several benchmarks and you compare it to models who were clearly behind on everything.
Give it some time.
Antigravity NEEDED to be game-changing. Without the stream of data that Claude, Codex, and Cursor enjoy there is little chance of getting an effective reinforcement learning loop. For the first time in its history, GOOG is at a meaningful data disadvantage, and apparently a cultural one as well.
Google literally giving everyone + student 18 month free subscription, those are source of cheap gemini + sonet,opus model that people selling/use with rotator proxy with thousands of account
they didn't lack the data
[deleted]
https://www.searchenginejournal.com/pichai-says-google-is-a-... (link to the actual podcast interview source within, this has a summary)
I don’t think this is necessarily true, did we all forget how much Google cared about alignment that their AI wasn’t able to render a white polar bear?
My view then was they are optimising the models for inference ability on their own hardware AND use cases, which is often speed and time to first token.
They've somehow seemed to end up with terrible compute shortages, which again is surprising given how good Google is at infra deployments AND have their own hardware. From rumors out there they are turning down enterprise deals for Gemini because they don't have the compute.
The problem is they're falling further and further behind on frontier class on coding especially, and since I wrote that article it's got even worse with open weights models undercutting them on price AND intelligence.
So my guess is that Google will continue having compute shortages until the Gemini enshittification starts.
Let's assume Google serves AI overviews on every SERP (they don't) and don't cache them (they do, afiak).
And let's assume that each AI overview is 2000 tokens (blended input/output), that's 500T tokens a month.
It's rumoured that anthropic is serving somewhere close to 10Q tokens a month.
Now it may be that AI overviews uses vastly more tokens than that per search, but I doubt it based on speed to render the overview.
My very rough napkin math on this is that maybe AI overviews is consuming 100T tokens/month max (after adjusting for caching and SERPs that don't have them), which would be 1% of Anthropic token volume.
Speed as a differentiator has always been Google's thing. They (used to?) show the microseconds it took to query & rank web-scale search results. Chrome, notoriously, focused on speed at the expense of resource use. The very many efforts to efficiently speed up Android & its runtime since its inception, and so on...
> their big model underperforms chatgpt 5.6
Possible but TFA claims:
We have started our most ambitious pre-training run yet, for Gemini 4 ...It would be a shame if they cannot beat Kimi K3 or Qwen3.8 Max, both of which are claimed to be Fable-like. If that is true, it will be [or would be] the first time a major American lab falls behind a Chinese competitor.
China can keep up because it's cheaper to run a frontier lab there. They also have more researchers and a stronger cultural inclination for this sort of thing. And I guess the business case in China doesn't have to work as well as it does in the US.
Not sure if this is what you meant, but their training runs are significantly cheaper. This was one of the big shockers from the Deepseek R1 paper. US foreign policy has helped to ensure that the Chinese are compute constrained, so they literally cannot buy the most expensive and powerful training rigs.
This has led to a steady drumbeat of innovations which are not revolutionary on their own but stack together to make things much more efficient.
like it literally pennies
When it is cheaper, and the "lower quality" model is adequate for the task at hand.
Plenty of problems have a low(er) skill/intelligence floor, anyone who uses the dual-mode agent paradigm (plan, then act) figures out the second phase can be completed by a less capable model. Even when disregarding costs - speed is important here because the agent can rapidly iterate without human supervision, based on compiler errors, lint and test failures
Paywalled article, but the headline is basically all you need: https://www.bloomberg.com/news/articles/2026-07-16/google-ge...
https://huggingface.co/microsoft/bitnet-embedding-0.6b
It’s a small multilingual embedding model designed for things like search, RAG, and semantic similarity. It supports a fairly large context window and is designed to run efficiently on a CPU in a GPU starved world.
The interesting part is that it builds on BitNet, using ternary weights of -1, 0, and 1 instead of the usual floating-point weights. That should make indexing and searching large amounts of text much cheaper without giving up too much accuracy.
the integration on ecosystem is the bread are
Lest we forget, "Attention is All You Need" came from Google.
It also came directly from the university of Toronto, and the university of Toronto seeded all American frontier labs (including Grok (why do you think they could start so fast))
I do not think OpenAI or Anthropic are actively chasing margins - though, Anthropic is supposed to be profitable on some form of non-GAAP accounting...
I suspect Google isn't really interested in seeing how far it can get dragged into a race of selling dollars for $0.25, and is more interested to see if it can stay in the race selling $0.50 for a dollar - when everyone else is losing or barely breaking even.
Maybe they don't want to price war with the other labs so they can comfortably maintain healthy margins on selling them compute?
it's their pro that isn't awesome at all. in fact, their pro kinda suck now that everyone else woke up.
Porting CUDA-based research, debugging, and overall experimentation speed is likely slower.
The GPU is still king for training.
Yes, subs like codex are heavily subsidized. But API billing has massive margins and that's what enterprises pay.
https://www.bloomberg.com/news/articles/2025-12-21/openai-se...
As for Anthropic, the rumors I remember seeing for their API margins were more like 85-90%, but I don't have a reference at hand for those. But once you know the API is wildly profitable and the subscriptions are roughly break-even and not even a big slice of their income, all of the investment makes a lot more sense.
OpenAI's new Mac app doesn't even have a normal "Chat" option now. OpenAI might be chasing coding and b2b sales more now that they realise very few regular consumers pay for subscriptions.
I think you overestimate long-term relevance of popularity among code monkeys.
I record whether the answers are correct, and the generation stats (costs, latencies, tokens used, etc.).
I have no idea why the Gemini models do so well.
I have recently added new tests, whose sole purpose was to find some cases on which Gemini 3 Flash fails (I don't like cherry-picking models or tests, but I also find it strange Gemini Flash models leading in accuracy). I made a more complex coding/tool-usage test, that I expected it to fail, it did fail it once locally in my debug tests, but when I finalized the test and ran the entire testing suite for all models, somehow Gemini 3 Flash still got it right...
Gemini models are REALLY intelligent (and they are actually my favorite model to use via the chat app to ask questions), but they somehow fail in real-word coding tasks where they have to modify files, check results, debug, etc.
My tests harness provides a lot of mock data, and limits the number of actions a model can choose from. I am starting to think that maybe the models are not bad, just that the coding harness are not optimized for those type of models, and Google doesn't really provide their own "Codex".
So yes, Gemini models are at the top, even if I actually (not proud of it) tried to make tests that actually favour other coding-focused models.
Also, would be great if you could add GPT 5.6 Sol XHigh and Fable 5 High as well, just to see if at least those beat Gemini which is currently your #1.
This does not match any lived experience or developer experience.
It'd be helpful to know _what_ you're testing and break that out by dimension. You mention randomly selected questions. How does that work?
With n=22 and binary pass/fail, the 95% confidence interval on a pass rate spans roughly (+-)15-20 percentage points. There's just not enough data ironically, for this leaderboard to mean anything. Ranks #5 through #25 are statistically indistinguishable
GLM defaults to max effort btw
https://docs.together.ai/docs/glm-5.2-quickstart#reasoning-e...
As far as I can tell it's slightly better than GLM 5.2.
[dead]
Ridiculous statement, pretty much all LLM tooling is model agnostic.
Disheartening, but not surprising: the comparison would not be very flattering for Google.
Light. Lite is product marketing seepage.
[dead]
The GCP team wants their slice, the other team wants some otjer slice, and so on. Everyone wants some crap for their promotion package.
It's no wonder Meta has shit the bed even worse.
It's also why Google still releases actually decent, useful models despite the product being such a hilarious mess. A lot of the time Gemini models have actually been better as production LLMs as part of LLM-based production applications than OpenAI and Anthropic models when it comes to the complete cost:quality:latency:adherence picture. And they still are. We have products in prod that use Gemini because they're better than any other model at the specific task. But we wouldn't dare use it for anything coding related, or even just as productivity tool to rely on, because as a consumer product it's a joke.
I got a Google One plan for Gemini, but it came bundled with YT Premium lite, and that somehow made it impossible to renew YT Premium for 30 days. I suspect different teams stealing customers from each other.
Like when you try to give Google money they try to squeeze you as much as possible.
At the same time you can get 5 time more limits for free just by registering 10 free Google accounts.
Google subscriptions are one big mess.
Haven't done any serious coding work with the flash models though — but I'm seeing more and more HN comments from people who seem to have picked it up for that in the last couple of months.
When they swapped Google Assistant for Gemini as the default voice provider in Android Auto it was so annoying. My wife's non-work space account can get Gemini to do the normal things like play music and what not, but my Workspace one can't do much of anything at all. I can talk about nearly any random topic with it, but getting it to change the playlist, nah, can't help you there.
It's no surprise to me to see them fumble actually supporting a lot of the consumer features of Gemini into Workspace.
[deleted]
They're not even benchmarking against other models now, just against themselves - which tells you everything you need to know.
There is absolutely no loyalty when it comes to coding. Nothing could be more common than people threating to jump ship whenever another frontier or open source model comes.
Google is clearly able to keep growing their free and consumer and small business use cases. Unlike corporate coding, we actually have evidence that solo and small businesses can actually see productivity gains.
Anthropic and OpenAI need to stay dancing like mad, because it's their source revenue which underpins their investments.
Why does Google need to shove something at the top at the same desperate cadence? Other than "recursive self improvement leads to AGI" it seems perfectly fine if they push out something dramatically better every year and half.
You should be happy for them.
This is still very early days. Who is "on top" has flipped back and forth many times already. The next frontier model release (from whomever) will change things again.
Claude Code has largely won individual developer mindshare and has been on top ever since it came out. The benchmarks change, but almost nobody opts to use anything other than Claude IME when I ask them. Enterprise is more competitive since they care about costs and other things, but developers leaning towards Claude puts a thumb on the scales there.
The product doesn't have much lock in, so it is possible to dislodge Claude, and Anthropic could (and some may argue is likely to) just shoot themselves in the foot again and again and again, but Google has never been particularly good at enterprise sales, and they have never actually been at the frontier of intelligence.
I think Google's incentives have mostly about building models for their products, which makes them focus more on the cheap end, and while they need that, it feels like the Innovator's Dilemma is biting them here.
I own a lot of Google stock from working there in the past and have been quite happy about their trajectory up until the last 6 months, but I am getting pretty antsy about their AI story these days.
Claude Code's success is not due to the agent but because the model is considered the best for programming and is very heavily subsidized, compared to pay as you go API prices. Consumers and Enterprise are not really locked in and will go where it makes the most sense.
I think they have almost no loyalty by actual developers.
At this rate, if Google has a flagship model, you're better off plugging it into a competitor's tooling than hope Google figures out how to use it.
They are not allowing me to hit their endpoints which agy hits - it's frustrating . i tried to hack it with gemini itself. what i love about gemini is it's so encouraging and ready to help you - even against the agy client : ) .
Even though im so frustrated with this - i still love Gemini for some reason ! Most encouraging model in the world!
And the real numbers could be better for Anthropic. It's feasible Opus models are actually cheaper to serve than GLM 5.2 because Anthropic have optimized the hell out of inference.
I guess that’s the big question, will people pay a big margin long term to use their end products / models or will AI tokens be commoditized by many competing players. For coding if I had to pay API costs I’d switch in a heartbeat, enterprise maybe more reluctant?
anthropic probably has more customers that use more of their sub, but for open ai where a lot of their subs are consumers through chatgpt.com, they have a lot of free money to work with there
Google's biggest and most important customer for all this AI stuff is google. Do they actually want other customers, or is having other people use their AI just an annoyance at this point, where we use up compute that they'd rather use internally...
And yet millions of people around the world are using their stuff, very often using only their stuff.
Likewise. This seems like a common feel. I have at least spent $4000 and likely a lot more on Gemini API because I really wanted them to win. I gave up.
Why do you care? Why would you spend your own money to a multi trillion dollar company so that they win their own "war" against another multi trillion dollar company?
Please don't get me wrong, I know the question can seem a bit negative, I am really just curious.
Although, I think saying 'wanted them to win' was not accurate. More like, I stuck with them hoping it will get better, and it did get better in many ways, coding was not one of them.
That being said, controlling android and apple mobile devices is kind of a big deal. And their video models are still top notch.
This is what happens when you put a McKinsey consultant in the role of CEO of an organization where product managers run the asylum instead of engineers.
* Abruptly ban me and all users from a Google Workspace for no reason whatsoever.
* Abruptly shut down a service.
I'll never give Google my direct money.
It’s so silly that individual people still use their shit. Corporations, I understand - they always choose the most mediocre stacks and tools by default. But why people choose to bring the mediocrity of Google into their lives is beyond me.
Because life gets simpler and better if you are just using what is available on your phone and in your browser instead of constantly chasing current HN darling.
Needless to say 3.5 was a disappointment. Curious to see 3.6.
HN lives in a bubble.
I have German/Italian/Polish clients virtually all use Gemini and NotebookLM. Talking insurance, banking, consulting, legal.
The real world doesn't look at pointless benchmarks on writing react tailwind crap, they are already google suite users, get the tools, test them and adopt them, end of story.
It's going to be like with angular, never mentioned on the net, widely used in the real world.
[deleted]
But hey, Gemini makes shittier job than Fable at producing my shitty "app" no one cares about. It is so over for Google!
That sounds awful.
For those of us who don't follow the AI hype cycle, what does that have to do with the topic of this thread: Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber?
What is Google's recently released AI coding product?
That is, given a lot of users' contexts, the way they can and the way they want to use these models are increasingly disjoint.
They're easy enough to skip - click the little "-" icon and you'll collapse the entire sub-thread.
It's a decent heuristic because the better models generate better pelicans. That's all. Nobody sane is going to make a bet on a model based on a pelican. But it's cool, it's tradition by now, and it's a semblance of a good first impression for new models.
But you bet my ass I check everytime to have a look at see how that pelican looks, it's just a fun check and also interesting to see the cost/results.
thank you.
I feel like the human brain massively overweights negative feedback over positive, and that's even after accounting for the fact that internet discourse tends to mostly surface negative comments (whereas the enjoyers stay silent). By default I have to try hard not to take it personally whenever someone leaves a negative comment about my work.
So just doing my bit to say I appreciate your commentary on so much of the fast-evolving AI landscape. Helps me orient :)
Consensus has an unfortunate habit of beginning that way.
[dead]
All of the pelicans so far have had really weird flaws / quirks so I am always a little interested to see how well these models perform at this task, since I've seen all the past pelicans and have some anchoring.
Seeing a truly flawless pelican would tell me that the model has true visual reasoning capabilities as well as good taste.
I agree to rednb that at this point it feels like rather obvious brand building, but also, I agree with you that some value is in it.
It does not feel all that authentic though, and it's good to react allergically to lack of authenticity. Bad for a lot of business models, but good for humanity.
Not only does it give you a super easy-to-grok understanding of the model quality just by looking at the image, but when you compare tokens and costs (both input and output), you really get a good, simple COST x QUALITY evaluation across models.
Simon explains it well: https://simonwillison.net/2026/Jul/16/kimi-k3/#what-can-we-l...
Simon, you should put up a summary table page that you update after every release.
But why is this an indication of literally anything else?
The 3.6 Flash pelican is just about the best I've seen.
[deleted]
[dead]
[deleted]
Sponsored blogs and paid newsletters are after all, notoriously poor at subsisting on silence :)
The fish and the cap where always added when I asked an llm to improve it's first attempt.
This continues the trend in LLM progress of better=more stuff
Edit: I wonder if this is a function of the reasoning training, where more tokens/ stuff is rewarded.
Does this say anything about the model? I meant the underlying attention/pattern it took for Gemini 3.6 Flash to create this SVG.
[deleted]
[dead]
That being said with any open model we of course do know the total cost (or estimate)
EDIT: It less less verbose in final output though, but it reasons more.
I assume the optimization comes when you have long-running tasks with many tool calls, and by reasoning more, it reduces the number of tool calls needed.
Given the extremely competitive releases of GLM 5.2 and DeepSeek V4 (both pro and flash), I don't think there'll be appetite for it.
i guess we'll use 3.0 flash but thats going to get replaced too right ?
these flash lite models aren't very reliable or consistent
—Ah, got it, it knows more about recent times.
[deleted]
[deleted]
You can get decent open-weight models now. That's not difficult. The difficulty is 1) running them and 2) compliance.
My company runs Claude on GCP's Vertex AI solution. We're in the US healthcare IT space, so the models need to be from somewhere that American healthcare agencies and companies have traditionally been okay with sourcing code from - which means the US, Canada, and maybe Europe. The stuff that handles PHI/PII must be in the US. The expense of hosting is more of a PITA than most customers want to go through this early in the technology's lifecycle, and intelligence gains are simply a matter of degree for most business tasks.
In theory, we could find some open-weight model (likely from China) for our development agentic work and host it anywhere you can host AI models. We don't, though, and I think Google, OpenAI/Microsoft, and Anthropic see that as the core of their business.
[dead]
Sure.
And how does that make your day better? I know it does not improve my work in any way shape or form.
I'll take a better coding model that's not multi-modal any time.
If I need an LLM to do images or sound, I'd rather use a dedicated one instead of a jack-of-all-trades-master-of-none model.
Or sometimes I will have tables, charts, or even screenshots of text that I would otherwise have to have another step to OCR or type out.
Multimodal saves me time on a regular basis. Not sure it’s a game changer, but just lets me communicate with the model in all sorts of ways that would be harder otherwise.
https://artificialanalysis.ai/#intelligence-comparison-tabs
Differences in token "density" are accounted for by pricing per task
but the implementation will be up to your provider and harness, for deepseek, they expose some numbers: https://api-docs.deepseek.com/guides/kv_cache/ and Anthropic has a list of actions invalidating your cache: https://platform.claude.com/docs/en/build-with-claude/prompt...
Basically, you avoid anything dynamic: model change, tool change, etc it's also important that your system prompt or main prompt doesn't have non-static data like the date/time/place or someone's name (the person you interact with in a chatbot for example). That should be left to tool call or search.
All of the models, you need to have a consistent input to get the cache hit. So if you are chatting with a document, and change the system prompt, it will be a cache miss, even if the rest of the items are all the same. If you even pass in the document in not the same order as the prompts, it will be a cache miss. Or if you add tool calls or structured outputs, it will be a cache miss. (Since those generally go at the beginning of the prompt call, not at the end.)
Most of the time when reading documents from URLs directly it will never cache. (Need to typically pass in the bytes directly, or use the provider document store index.)
Gemini has a 4096 minimum token size with the 3 version models before even getting a cache hit. OpenAI it is lower (1024), and is automatic, but only happens in increments of 124. Anthropic can also get cache hits at 1024 tokens, but you need to explicit ask for it (and pay extra).
Caching by default typically lives for 5 minutes since the last cache hit across providers. But some of them you can ask for longer. AWS for Anthropic models can be tricky with multiple endpoint routing, so can get cache misses if it happens to route to a different endpoint.
[deleted]
Features stay in Beta for ages, whatever that actually means, and released ones get deprecated things fast.
Where some of the competitions treats deprecating entire services as "let’s not put it on your frontpage, put deprecation notices all over the doc, and politely ask new users not to start new project with them".
These days no company even has completion models where one controls the text input fully. Worthless.
google's inability or unwillingness to provide stable timelines for model deprecation makes it risky to build complex workflows using their models
For my use case, `gemini-3.1-flash-lite` is ~20% higher accuracy than the next best model of comparable cost (considering both proprietary and open-weight alternatives)
They are not anywhere close according to pretty much every benchmark (even v4-flash is considerably ahead and its way cheaper than flash-lite). Maybe tuning prompts/tools/etc. might be useful?
I presume you can't use deepseek?
You can also just write code like you did a year or two ago.
[dead]
If you're using it for other purposes, then I give you permission to ignore my comment; there's no reason to descend into name calling.
we are switching to Deepseek.
[dead]
It might be overkill features-wise, but there's a free tier and it likely won't be left for dead anytime soon.
[dead]
And how is that different from their competitors exactly?
[dead]
I do think it's still also simultaneously true that they have an actual problem with competing with current frontier progress. It's just that has gone from an existential threat to something they are willing to defer addressing because they see the long game for them sitting at the smaller end.
You don't have to self-host the open-weights model, you just need to be able to source it from multiple providers.
Using the closed vendor models maybe made sense when open-weight models lagged so far behind, but that time is now gone.
Have you considered moving to open source / Chinese models?
If gemini-2.5-flash-lite is good enough for your application, you will find even lower cost options with better performance outside of the Google ecosystem.
artificialanalysis.ai has it going from 165 tps -> 304 tps. openrouter.ai needs more data but it has it going from ~100 tps -> ~150 tps, though at peak 3.5 has reached 156tps.
(I actually use a mix of both for some offline projects, nothing serious.)
'Antigravity' which is an agent-first editor layout where you don't see your code and just prompt it.
'Antigravity IDE' which uses the Windsurf/VS Code editor, which is still what I primarily use in my day-to-day.
I hope they never retire the IDE, I don't think I can get used to prompting an AI Agent without being able to see my code to help workout what needs to be done.
Being able to ask questions to small open models seems.... obviously useful?
Use the right tool for the job. It's like asking why a screwdriver isn't good at sawing wood, or calling C a terrible language because it's hard to make CRUD apps with it.
Rankings for text are here https://arena.ai/leaderboard/text
For comparison of Free Tiers: - Gemini serves 3.6-flash (rank 12) - ChatGPT serves 5.5-Instant (rank 23) - Claude serves Sonnet 5 (rank 27) - Meta AI serves muse-spark-1.1 (rank 5)
While Meta AI serves the better ranked model, it doesn't end up working that well for other things. For example, if I ask "help me buy a new raincoat", it ends up suggesting a Cambodian website, whereas Google is well integrated with Google shopping. It doesn't have the same integration with GApps outside of Gmail/Calendar. A few other email connectors are available.
Claude has one of the best interfaces with connectors, skills, and plugins galore, but the model and limits are restrictive on the free tier.
Gemini, as far as I know, I've never hit a rate limit on Flash.
I believe Gemini is going to gain market share through the free tier funnel while serving models as cost-effectively as possible. People are going to use Gemini because they use GApps and Google.
ChatGPT and Anthropic are going to be competing for the API/Business users, but for everyone else they are going have to become Google before Google becomes them.
(it lets you search youtube by transcripts)
Seriously?
[1]artificialanalysis.ai
I find that quite staggering. GLM is open weights
[deleted]
Anyway, given that both Gemini and OpenAI have 3 sizes of models, one would think Google compares their medium size to OpenAIs.
Why do you care that much? Just say it. You're letting an imaginary score determine if you should express your suggestion for improvement. That's a bit wild.
I prefaced it with that to simply say I knew it'd be an unpopular opinion.
It gets worse from there. Gemini is terrible at agentic coding, primarily because Agy is terrible. I noticed Google updated Agy with this release, so perhaps that's finally going in a good direction. I'll have to test it. Thus far, my experience in countless experiments has been Gemini models being 2x faster yet with less depth in their solutions and a lot more going off track.
I very rarely have to stop Codex or Claude Code sessions because they're doing something random and unexpected (or not asked for). I genuinely think Gemini models are brilliant but virtually useless in agentic scenarios in my experience of the last few years (2.5, 3, 3.1, 3.5).
I should also note that Gemini web UI annoying resets to its lowest intelligence which feels scummy and Google is not transparent about what "extended thinking" really is. Past posts have pointed to "extended" being medium. Every other provider gives you the raw value (medium/high/etc).
So honestly, I feel Google would rather I don't use their models. They just want to get a little bit of mindshare and stay in the conversation. I had the Ultra plan and cancelled it once it was apparent they were not improving the agentic experience nor trying to compete.
PS: That opus-4.6 via agy works so much better points in another direction though!
https://www.wsj.com/tech/ai/chinas-xi-touts-open-source-ai-a...
So who even knows.
funny because some people downvoted me believed that there is no relation between knowledge cut off date and real world events. that's not how it works!
Damnit, I usually don't jump to LLM speech patterns, but this opening had me thinking you were a bot. But after checking your profile, I think you pass as human. I wonder when will be the time, this does not work anymore for me. (Creation date is a strong hint, but abandoned accounts can be hijacked)
"10 Quadrillion tokens a month means: 333 Trillion tokens per day and 3.85 Billion tokens generated/processed every single second, 24/7."
"At an incredibly cheap, subsidized infrastructure cost of $1 per million tokens, serving 10 Quadrillion tokens would cost Anthropic $10 Billion per month ($120 Billion a year) just in inference compute."
It also had this to say about how google's AI overview works: "Google doesn't just feed the LLM your 5-word search query. The system scrapes the top 10–20 web results, feeds thousands of words (tens of thousands of tokens of context) into the model, processes it, and then outputs the result."
Oh, and it does all of that in less than two seconds. Honestly, whatever Google is doing with its infrastructure is so far ahead of everyone else, I can't believe you fell for such an obvious lie.
There are also extremely obvious holes in your comment:
>Let's assume Google serves AI overviews on every SERP (they don't) and don't cache them (they do, afiak).
Try it out for yourself. Add a few random letters or punctuation. They cache nothing.
They definitely cache results - I've searched and re-searched an identical query back to back a few times and seen identical results from overview. They are definitely throwing a stupid amount of compute towards these ai results nobody is paying for - changing punctuation and stuff does get you a different response - but they're not doing no caching.
Certainly what they're doing with their infrastructure is impressive but it's not super meaningful at the end of the day for a for profit company to be really impressively good at burning tens of billion dollars on a service nobody pays for while the same tech from their competitors is quickly becoming one of the largest spend categories for many software engineering teams
That said, Anthropic has supposedly crossed over into profitability and made $1 Billion in profit so far this year, in the lead up to their IPO. Being profitable sounds good for launching on the stock market! But as a customer, that's noticeable in the downtime due to lack of compute, and only getting 50% access to Fable.
OpenAI might not be profitable, but they've got so much compute access that they've been able to give their customers full access to Sol, and as a result they've almost doubled their Codex subscriber base in the last two weeks (6 million on July 12, 10 million on July 21 - that would be an extra $1-$10 Billion in Annual Recurring Revenue that they've gained in just these 2 weeks). Doing the unprofitable thing in the short term can result in outsized rewards in the long term.
> would be great if you could add GPT 5.6 Sol XHigh and Fable 5 High as well
I would like too, but I avoided them for several reasons:
1) Cost - this is a hobby project, those models would cost tens of dollars for each benchmark run, multiply this by tens or hundreds of models and ...
2) Time - the high models are already taking a really long answer to respond (5-10minutes per question). I run each question with 3 repeats (run the same test three times), so it would take 30 minutes per test. If I change my tests, methodology, or add a new test, it would take a really long time to run the benchmark. Also, I like having results immediately when a new model is released, now I can post within 30 minutes of a model's release the benchmark results.
3) High reasoning usually does WORSE on most tests - if you look at the leaderboard, it's sometimes counter-intuitive, but models with high or max reasoning usually do worse than medium and low. This is because the questions are quite targeted/direct, and the models overthink the question and miss the solution. Or the long thinking context makes them perform poorly. The generation tasks (SVGs/HTML animation) are usually better with longer reasoning, but short code fixes, trivia questions, puzzles, etc. are answered by low/med reasoning with more accuracy in general
Also, Fable is borderline un-testable, it refuses to answer many questions, so it scores poorly anyway.
Gemini scores 21/22 because it answers all tests, and it does them correctly, consistently. The only failed test is I think because it miscounted the lines in a file, when responding on which line the bug was in a code snippet.
There is some short info about the methodology here: https://aibenchy.com/methodology/
> GPT-5.6 Sol on Low beats Fable Medium by 10% > Fable is number 20
Fable loses a lot of points because it often refuses to answer questions. Asking a basic tool-usage challenge, Fable responded with refusal: "This request triggered restrictions on violative cyber content and was blocked under Anthropic's Usage Policy. To learn more, see https://platform.claude.com/docs/en/build-with-claude/refusa...." Even in practice, you ask Fable something trivial, and it refuses to respond. I think the score accurately represents how the model is behaving in real-world usage.
> Gemini-3.6 Flash then beats them both Gemini models are the most intelligent overall. The tasks are not coding-only. Gemini excels in general knowledge and domain specific knowledge. Gemini models, even old ones, still top many charts on specific use-cases[0][1]. Depending on how you weigh those cases, the leaderboard order can vary quite drastically, as some models are very strong in some domains and weak in others.
> You mention randomly selected questions. How does that work? Randomly selected, means I have manually created the questions/challenges to span across various domains and agentic surfaces. Questions vary from coding tasks, tool usage, trivia questions, chess puzzles, car-wash-like challenges and more.
> With n=22 and binary pass/fail Each test is run 3 times, so in total we have 66 tasks. Also, apart from correct/wrong answer, the final score also includes the pass rate for each test (how many out of three attempts), how good the reasoning is (they have a hidden reasoning score where available) and other small factors. Also, some tests in some categories involve a series of tasks/requirements (i.e. implement this function, call it, do some processing on the result, combine the result with some built-in knowledge data, etc.).
I do agree that 22 tests isn't that much, and I'm slowly adding more, but even without the leaderboard part, the comparison feature is what's I think is most useful. You can see for the exact same tasks, which models do better, which do it faster, which cost less, etc.
> _what_ you're testing and break that out by dimension There is a category breakdown for the test results, so you can see and which sort of tasks models fail.
Everything aside, when you manually ask a model to test its capabilities, I don't think it takes many questions to realise how good/bad that model is. Sometimes one prompt is enough, you ask it to do something, and see how it reasons about it, how fast it does it, how efficient the steps are and how good the result is. Yes, the performance may vary across tasks, but I'm pretty sure if you did a blind test with a chatbot, you could easily realise how good the model is in just a few questions/tasks.
I think no benchmark is perfect, mine is far from it, but it's simply another different, independent data-point. Apart from that, I made this for myself, and I'm using it myself. I don't trust that all popular benchmarks are not in the training data, and I think many benchmark the wrong things which don't correlate to how I use the models day-to-day myself. I just made the results publicly available, in case any one else benefits from it. I've probably spent thousands in LLM costs, and probably more than 100 hour building this, without benefiting in any way from it (outside of the joy of building it and me using it personally to compare models); as long as models cost stays reasonable, I'll keep building it and test new models as soon as they are released.
[0]: https://x.com/browser_use/status/2079602472516264010/photo/1 [1]: https://artificialanalysis.ai/evaluations/mmmu-pro#mmmu-pro-...
If it weren't for the $10 GCP credit, I'd straight away cancel it. I don't see enough value in Gemini to justify the $20 subscription.
They might have success if they tried maybe a 12-15 dollar tier.
Build a whole new management tree - the current people all do a terrible job.
I have multiple anthropic and OpenAI max plans. For Gemini I just use my Cursor $200 a month plan (which also gives me the ability to try grok, conductor, etc)
[deleted]
When we work trial people, 100% of people ask for Claude rather than Codex or anything else.
When I talk to people at non-AI tech events everyone basically says they use Claude and have not tried an alternative.
Developers writ large are actually not that interested in trying multiple tools, they like customizing their chosen tool and tweaking it forever.
I think developers are as susceptible to brand marketing as everyone else. It's why almost everyone has a Macbook.
Basically you may choose to drink brand A water bottle, brand B water bottle or tap water. Oh and you might choose the glass water bottle if you use API/Fable.
It's not true, it's just a play for margin.
Where does that 90% figure come from?
I know some down to about 1% ($200 Max plan vs $20k in tokens per month)
Not every subscriber is a full time SWE. In fact most professional SWEs will be on enterprise plans and thus not get subscriptions at all.
I think it's very plausible that subscriptions are overall losing money. But we simply don't know.
(and if you are using per-token billing as an enterprise user without first maxing out a premium seat you are very silly)
And choosing one of them obviously makes life more complicated than not doing that.
> getting your ass profiled
Gmail knows nothing of my ass.
> let me know if the ads in your mailbox are helpfully tailored to your needs and preferences!
I couldn't, for I don't check them. I have enough life to live without degoogling crusade.
And outrage at the unimpressed is.. a curious standard for a sponsored blogger.
"He doesn't have a tin can in your face shouting from a soapbox that makes it hard to ignore."
He metaphorically does because people upvote his pelicans to the top and the ensuing comment threads are massive/bloated. Huge amounts of the readership of this website are lurkers who don't even know how to hide these giant posts. Look at how bloated this very thread is right now!
Also, a lot of people unironically are whining about him because of sour grapes. Pay them what Simon is likely making, give them as much mindshare/attention as Simon gets, and they wouldn't be so mad.
Anti-incumbency bias and anti-elitist attitudes are good actually.
Agreed, but scrolling past or hitting - were apparently off the table for the complainers. So flagging was yet another tool in their toolbelt and I bet a powerful one at that. If simonw's pelican posts routinely went dead from flagging, he would not make them. You know that, I know that.
> The pelican test continuing to be taken seriously is a great example of that kind of echo chamber.
You can try and support that argument if you like. But I would implore you to realize that it has been had many times recently and the other side does in fact find value and do not see it that way.
> He metaphorically does because people upvote his pelicans to the top and the ensuing comment threads are massive/bloated.
Users upvote the pelicans because they find it interesting. they arent paid trolls or simonw fanatics.
> Huge amounts of the readership of this website are lurkers who don't even know how to hide these giant posts.
They can learn... it's called hackernews. For those interested, that is what the [-] link is for above the comment. Use it and move on.
But, also, as said, you're an industry (and foss!) veteran, so I find it impossible to believe that you haven't had your fair share of baseless bullshit being thrown at you, and with that, you gaining a persona that will not be hit by that, because it clearly knows that it is in fact bullshit.
Unless of course it doesn't really know that with certainty.
As said, I would _love_ to give you the benefit of the doubt, because you might just have a stressful day or whatever, but content marketing is literally your whole thing by now. It is impossible for me to do that with a clean conscience.
Your blog front page currently opens with
> Earlier this month I hosted a fireside chat session at the AI Engineer World’s Fair with Cat Wu and Thariq Shihipar from Anthropic’s Claude Code team.
That is not what "some rando foss maintainer we are morally obligated to be soft with" does.
But I repeat myself.
I think very hard about the ethics of what I'm doing and how I can best use my "platform" (shudder again) in as constructive a way as possible.
[dead]
[dead]
> It does not feel all that authentic though, and it's good to react allergically to lack of authenticity. Bad for a lot of business models, but good for humanity.
I hope SimonW keeps them coming.
I get people burning out on the pelican SVG test alongside the rest of the AI burnout, but I guess for myself I'm just choosing to keep enjoying it while I still can.
Marketing, apparently, is a property of URLs rather than outcomes :)
If you’re constructing the prompt you don’t have to jam everything together you can arrange it appropriately.
[deleted]
I'd rather you do this than him. Chill out, please. The pelican pic is fine.
There's nothing shady here. The disclosure is front and center on his About page on his website.
He's not spamming you. It's one short link, sometimes a link to a first-impressions post. It's interesting and useful for me and the other commenters who keep upvoting his comments. Why are you so antagonistic?