Here's GPT-6 Luna pelicans: https://tools.simonwillison.net/markdown-svg-renderer?url=ht...
And GPT-6 Sol: https://tools.simonwillison.net/markdown-svg-renderer?url=ht...
Scroll to the bottom for the GPT-6 Sol max one: https://tools.simonwillison.net/markdown-svg-renderer?url=ht...
For comparison, here are the pelicans I got for GPT-6 Astra: https://tools.simonwillison.net/markdown-svg-renderer?url=ht... - I still like the Astra Max one best.
Here's a comparison grid showing all of the GPT-6 and GPT-5.6 pelicans at all effort levels: https://static.simonwillison.net/static/2026/gpt-6-and-5.6.h...
The grid is actually really interesting, because it shows that the 5.6 family default to brighter colors than the 6 family.
1/ Usage limits: downstream of input/output cost, but resets and obscure windows and odd 20x plan / 5x plan != 4x usage math throw a wrench into it. Winner right now is Codex by a mile, especially when you factor in the fact that ChatGPT usage (even 6 Astra Pro) is essentially unmetered on the 20x plan. Always a bummer when asking if I should see a doctor about a rash means I can't code as much. It's also is a godsend if you use an MCP like oracle to automate the process of calling the Pro model on particularly tough problems, giving better planning results or deeper code analysis without burning usage.
2/ Context window in the harness. Claude Code wins on this. There used to be a toml file workaround for Codex to extend the GPT context window to 1m, but this stopped working on the plans and only on per-token billing (ETA: noname120 pointed out this is no longer the case and it can be enabled again [1]). 252k is just not enough. Codex's compaction is very good, fwiw, but it happens so frequently that even a model as powerful as Astra sometimes loses the plot on long-running tasks.
3/ Ability to use the plan outside of the official harness. Codex wins. Anthropic does shit like bills requests as extra usage if it sees a hermes.md in a commit.
I've subscription hopped a bunch, and at times I've had both, but I keep coming back to Codex because it wins on 2/3.
ETA: apparently I haven't been Keeping Up With the Altmans and new 20x signups have been disabled for a few weeks. I am grandfathered in, which makes the comparison above pretty much moot.
- For general chat and web search, occasional image editing, small coding work, document review etc. ChatGPT Plus is basically limitless and “just works” since 5.6. I’ve yet to give it some task it cannot do.
- When given sensible instructions, it hardly annoys with weird phrasing, glazing, or annoying constructs.
- The apps are very good (ignoring the initially terrible Codex app)
It’s easily my best spent $23 a month.
Prices per 1M tokens Claude Opus 5.5 Claude Opus 5
Cache reads $0.20 $0.50
Input tokens $4 $5
Output tokens $20 $25
Cache writes $5 $6.25
ModelInput
Output
Price reduction
GPT‑6 Sol vs. GPT‑5.6 Sol
$4 → $2
$20 → $10
50% cheaper
GPT‑6 Luna vs. GPT‑5.6 Luna
$0.20 → $0.10
$1.20 → $0.50
50% cheaper
I actually preferred 5.6-Terra not because it is technically superior (it isn't) but because it had better instincts to NOT do this stuff.
PS - Speaking of better instincts, have they closed the UI-design gap at all? I keep a Claude subscription just because /design produces significantly higher quality UI design/UI feedback/UI refinement than anything I've seen from OpenAI.
Edit: Yes, it applies also to subscriptions, source https://x.com/thsottiaux/status/2102463847714247142
EDIT: this doesn't say anything about availability on either Azure or AWS. I'm assuming it will show up later, but it would be interesting if it didn't.
Design: https://image.non.io/78795662-8bfc-4e14-8d72-3738392aa6b3.we...
All 3 were given the same prompt to dynamically light these and to create the designs as a SPA with page transitions.
Astra: https://html.non.io/annui-astra
Sol: https://html.non.io/annui-sol
Luna: https://html.non.io/annui-luna
Luna gets the button wrong, and in the same way Grok/MiMo did. Looking into it more, it's because Luna actually searched my computer for similar builds, found the ones that I did for grok/mimo, and referenced their files. Astra is still the best by a significant margin in my eyes. Far more polish, better page transitions, effects that aren't overcooked and take into account the page. Better contrast.
Now GPT 6 Luna is even cheaper, and more intelligent, there is no going back... to SOL 5.6 for intelligent layer.
I was hoping for a serious Luna upgrade. It was already cheap enough. This feels more like a price reduction than an upgrade.
That said, if the new Luna is able to handle ultra mode and subagents v2 in codex cli, then at least that’s a win.
When 5.6 dropped I had no weekly limits and I could just drive my work with Sol xhigh and things were great. Once limits were back (and maybe token prices changed iirc) Sol was no longer usable (on Pro or business) unless I was ok with 4 prompts every 5 hours, so I had to switch to Terra medium/high. I've used Luna for some really dumb tasks like moving files, renaming variables and whatever other old-school refactors I've needed.
Then Astra dropped and it just uses so many tokens I've only prompted with it once. Now with GTP-6 Sol/Luna I'm not sure what's being said here but most importantly I'm wondering whether Luna 6 is a good replacement for Terra.
Has any other Terra user tried and knows more or less than answer to this?
All I want to know is how old is the model and how much does it cost. I can figure out which one I want to use based on that, assuming that newer models are always better.
Trying to convince us there is a difference between GPT-6-Sol and GPT-5.6-Terra or whatnot is ludicrous to the point of being insulting, especially when new models come out every week.
[0]: https://aibenchy.com/compare/openai-gpt-5-6-terra-high/opena...
> In the Coding Agent Index, Sol improves but Luna regresses: In OpenAI's Codex harness, GPT-6 Sol (max) scores 57 in the Artificial Analysis Coding Agent Index, up 2 points from GPT-5.6 Sol (max), with gains in Terminal-Bench 4.0 (43% vs 37%) and SWE-Atlas-QnA (58% vs 54%). At $2.99 per task it costs ~50% less than GPT-5.6 Sol (max) and sits on the Pareto frontier of Coding Agent Index vs Cost per Task. GPT-6 Luna (max) scores 41, down 2 points from GPT-5.6 Luna (max), with lower scores in SWE-Atlas-QnA (44% vs 49%) and DeepSWE v1.1 (64% vs 66%), at ~60% lower cost per task.
https://x.com/ArtificialAnlys/status/2102462962758033624
Given that they had to discontinue sales of the 20x Pro plan after the Astra release due to compute constraints, I wonder if 6 Sol & Luna are smaller vs their 5.6 counterparts?
GPT-6 Astra (low medium high xhigh max ultra)
GPT-6 Sol (low medium high xhigh max ultra)
GPT-6 Luna (low medium high xhigh max ultra)
And that's not even counting the GPT-5.x models: GPT-5.6 Sol (low medium high xhigh max ultra)
GPT-5.6 Luna (low medium high xhigh max ultra)
GPT-5.6 Terra (low medium high xhigh max ultra)
GPT-5.5 (low medium high xhigh max ultra)
And then there's a fast mode toggle for all of it, too.Not exactly a low-friction user experience!
Like are you supposed to just somehow intuit, “ah yeah, this task is definitely a GPT-6 Sol Medium task,” or something?
Is this just second nature for OpenAI employees? How are end users supposed to know how to optimally choose a model for a given task? Am I missing something completely here?
So essentially I was not able to get nowhere close to the accuracy of previous model and it was slower, and more expensive at the same time.
Now gpt-6-luna, has really competitive pricing and offers similar accuracy compared to gpt-4.1-mini fine tuned for my specific task. And fine tuned models are getting deprecated anyways, seems like a good time to move to gpt-6-luna.
At least it sounds good on paper, the the graphed results do give me pause as it seems the lower cost might come from a slightly nerfed base model combined with more thinking, going by the more erratic scoring curves and the lower no thinking baseline score. I’ll have to try it out but I really hope they haven’t nerfed Luna/Sol to make this price point possible!
Overall, I expect for most people think the winner of today was Anthropic. I personally am preferring Opus 5.5 at medium over GPT-6 Sol Max, in very very early tests. Similar price range, more capability.
But competiton is great, these are solid releases by OpenAI today.
Model update Input Output Reduction
------------------------- ------------- ------------- ---------
GPT-5.6 Sol → GPT-6 Sol $4 → $2 $20 → $10 50%
GPT-5.6 Luna → GPT-6 Luna $0.20 → $0.10 $1.20 → $0.50 50%Does anyone know how exactly these price differences for example between sol6 and sol5.6 translate to codex percentages? In theory it seems like for "high" on both it should result in ~3x more usage. If that is actually the case it would be huge! But all we see is % left and % changes while using and we really have no idea when or how those numbers are being calculated or when they change. So there is a 50% price reduction on API but who knows how the hell that translates to whatever price calculation is used on codex.
Which... fine, I'll take that.
Rarely have I seen such hogwash. It seems to be a mix of virtue signalling and trying to push the perspective that being "90th percentile" (on what exactly?) requires extensive AI use. You are telling me you expect each researcher to generate USD 7000/d or USD 140k/m in AI cost? Or is that a way to abuse tax laws in some way so they can claim their own payments for tokens as expenditure on the other side of the ledger?
For example, I had Fable review Astra’s output yesterday, and it found some issues and fixed them. Passing the fixes back, Astra then uncovered additional issues with Fable’s fixes (and yes, this will go on ad infinitum if you let it, but these were “real” issues).
It seems the big story here is the reduced Luna pricing. It’s a fantastic model that can handle most automation needs (though I still use the big models for day-to-day development).
* Prompt caching dashboard: https://platform.openai.com/usage?usage_section=prompt-cachi...
* Adjust reasoning effort and tool availability without breaking cache
Wierd!!
https://community.openai.com/t/experimental-context-manageme...
I would also like to point out that it was quite predictable that Terra got discontinued, it didn’t make sense to have it when both Sol and Luna overlapped it.
Lunas insane discount is a game changer, OpenAI knows what they are doing here. Luna at max reasoning effort, even though its not optimal for long conversations, its incredibly intelligent while dirty cheap. Its not even competition anymore.
Whats even crazier is that I’ve underestimated how good Luna actually is. I’ve seen colleges create fantastic things with just Luna medium. This basically means you never have to think about your Codex usage anymore. You can run all day and not
have to worry about your 5h or weekly usage limit. To me, the discounts OpenAI is offering with Sol and Luna is truly a new milestone.
GPT-6 Luna now is 50% cheaper, which makes it have one of the best intelligence per cost ratios.
GPT-6 Sol is smarter, but seems to reason 2x more than GPT-5.6, which makes it 2x slow3r and 25% more expensive in practice.
[0]: https://aibenchy.com/compare/openai-gpt-6-sol-high/openai-gp...
Now it is revenge time and OpenAI kind of is trolling Anthropic by simply going into a price war with impressive performance.
OpenAI is doing a decent job this year after they recovered. Anthropic needs to offer more payment options and be clear about token usage. The warnings I got when switching to Fable 5.1 felt like a thread. I bet more and more on OpenAI since I don’t feel robbed by them.
These models are significantly cutting down the token costs by almost 40-50% as compared to their predecessors. This is exactly what people need - cutting edge intelligence at half the cost.
> GPT-5.6-Sol is retiring. This conversation will automatically switch to GPT-6-Sol
I don't recall OAI retiring a model so early lol. Similar arch?
It seems there's a bug, shipped together with the flag that enables the new models, that doesn't allow Codex to run properly in the WSL2 sandbox.
How do you decide what to pick? I mean, I do Platform work on a large monorepo with many different interconnected services, and so I always want the implementation to be "correct".
Incredible.
Does anyone care about code quality anymore?
I guess the chinese competition spooked them.
Have the frontier labs stopped trying to increase context window size?
> On FrontierCode, which evaluates whether coding agents produce changes ready to merge into real codebases, GPT‑6 Sol improves substantially over GPT‑5.6 Sol, and is able to match Claude Fable 5.1 xhigh at much lower cost.
I continue to appreciate OpenAI's attempt at some honesty here, showing that they are capable enough and have skilled engineers to a point where they can recognize that slop is hated for good reason, and that there is a real issue. Compare this to anthropic, where e.g. in the Opus 5.5 announcement[1] one of the first points on the page is
> One tester completed a 680,000-line code migration in less than a day—work that would have taken an engineering team weeks. It’s good at finding and fixing inefficiencies in software: when we asked it to cut load times across every page of a web app, Opus 5.5 succeeded 39 of 40 times, while Opus 5 made smaller improvements that also altered the app’s behavior. A different tester had several Claude models build a game from a single prompt; Opus 5.5 scored higher than any other model on the strength of its graphics and polish.
This is the kind of shit that is the very reason why I stick to OpenAI and deepseek. OpenAI is simply more honest and reasonable about their models' capabilities, while delivering models that still have solid value.
Notice how the OpenAI announcement doesn't make use of anecdotes.
OpenAI seems really competitive in most areas, and extremely competitive on cost, but still behind on coding.
[deleted]
I kind of hated Astra for it's poor instruction following and stopping all the time plus bad code quality. It somehow feels a bit like some of the popular open models but with a lot more knowledge or peek capability. But it doesn't reach peek that often
The models below are competing on value for money.
In service development coding, a top-level model is not required.
I hope this kind of competition continues.
Has anyone else noticed this?
[deleted]
[deleted]
From around GPT 4 results got "Good enough"...I generally try to explain what problem I'm trying to solve, set limitations and boundaries, tell it to ask me questions, have it write up a plan with steps then we take one step at a time.
These new models are starting to feel like iPhone releases where the improvements / feature set feels incremental.
Same on the Claude side which I use for work
[deleted]
[deleted]
I wouldn't be curious to sign up to codex whatsoever these days
These token reset shenanigans are insane
Me: "Can you check this thing?"
Sol: "Of course!"
Me: "Do so then!"
Sol: "Ok, I checked it."
Me: "Aaaaaand?..."
Sol: "I found some verify significant things."
Me: "List them! Actually, you know what, let me just go back to gpt-5.6 this is ridiculous."
OpenAI is promising "the Sun, the Moon, and the Stars". The spirit of P.T. Barnum is doubtless looking on with jaw dropped at what is beyond doubt one of the greatest demonstrations of chutzpah, by some of the greatest hucksters, in the history of the human race.
As for any cost based argument, it is immediately invalid because the cost is something that OpenAI fully controls and manipulates.
;)
[dead]
[dead]
[dead]
[dead]
[dead]
[dead]
[dead]
[dead]
[dead]
[dead]
[dead]
[dead]
[dead]
[dead]
[dead]
[dead]
On the other hand, you have previously written: I'll gladly admit I think what these companies are doing is unethical, and I'm sure that biases my thinking toward skepticism. [1]
You have now posted accusations/assumptions of astroturfing and manipulation at least 15 times, without ever providing any evidence. This is in breach of the guidelines, because comments like this poison discussions far more than the comments they're complaining about.
We – of course – want all comments and posts on HN to be authentic. HN is only a place where anyone wants to participate because since the beginning, we've had software mechanisms and moderation practices that detect and weed out inauthentic commenting and voting. We're identifying and dealing with it every day, continually improving the software to detect and remove it. Most of that happens quietly and efficiently in the background without anyone having to see it. When users see evidence of manipulation and report it to us via email, we happily and thoroughly investigate it.
Most of the time, what we find is simply that people are authentically excited and passionate about the topic, which is what is happening here. I understand it can be hard to accept that if you're skeptical about the topic.
It's fine to be skeptical about the topic and you're welcome to express your skeptical views on the topic. People do that every day on HN, about AI-related topics and countless others. Healthy debate is what we're here for.
But you can't keep poisoning HN, by (1) continually posting these unfounded claims, then (2) when users and moderators simply uphold the guidelines, staging a protest by demanding your account be deleted. This is not what people do when they care about a forum's health.
[dead]
[dead]
[dead]
Cached Read: ~6,500M
Input: ~150M
Output: ~20M
Approx $40 worth of usage across DeepSeek V4 Flash + MuseSpark Contributor 1.3. And a bit of both the GLM models. This is covered in a $10 subscription.
If I were to use Luna's API pricing:
$0.02 x 6,500 = $130
$0.20 x 150 = $30
$1.20 x 20 = $24
So $184. And this is assuming smaller coding sessions (<272K) beyond which Luna pricing doubles.
--
Cost wise, these models are nice for small stuff. Translations etc. Any model that does not provide multiple Mtoks of cached reads per cent is not very useful to me for coding workflows.
Forget DS. I asked MiMo 2.6 yesterday to explain ML/LLMs to me succinctly and the pointed it at Karpathy's micrograd code. It produced a C implementation called `xor_mlp`, a tiny model that learnt how `xor` worked. I then asked it to produce a model that can play tictactoe without losing (mostly). It did. It supervised the training process and produced a compiled version with multiple switches. The pi-dev session is still running, so here are actual stats
↑45k ↓35k R1.0M CH99.4% $0.019 4.2%/1.0M (auto) - (opencode-go) mimo-v2.6-flash • high
And here is Luna on the same workflow (I had to poke and prod a bit to get what I wanted):
↑141 ↓34k R1.0M W43k CH95.3% $0.072 4.2%/1.1M (auto) (opencode-go) gpt-5.6-luna • high
I expect similar results from DS41F/MS13. Closer to MiMo costs than Luna.
So the "significantly cheaper" thing may not really hold, more so when Luna has to actually read my codebase to do the stuff that I want rather than rely on world knowledge. The 8-10x cache read cost differential itself will kill the token budget.
I don't think you can guess more precisely than an order of magnitude from trying each once on one task.
"Artificial Analysis Intelligence Index combines performance across 10 evaluations: AA-Briefcase v1.1, GDPval-AA v2.1, AutomationBench-AA, Terminal-Bench 4.0, SciCode, Humanity's Last Exam, GDP.pdf, CritPt, AA-Omniscience, and AA-LCR v1.1."
Not saying it doesnt have any value but it's probably irrelevant if you use these AIs for a specific use case. Like for example Humanity Last Exam tests general knowledge, which is not very useful for coding.
It's best to go to the specific coding benchmarks and compare there.
My cost is (I use nous as provider)
DeepSeek v4-flash-0731 • Your cost: $0.56
DeepSeek v4.1-flash • Your cost: $1.22
GPT-6 Luna • Your cost: $4.22
My usage is heavy on the cache. Apparently v4.1 flash uses 1.75 times as many tokens so still cheaper.
Frankly, I have no idea what people do with Opus/Fable etc. I don't think anything I do needs something that charges $50/M for output tokens.
I need aggressive cache read pricing with full prompt_cache_key support to have a model be financially viable for our workload. Right now Meta Muse 1.3 Contributor is the only one that makes sense--but we are starting Evals on the new MiMo 2.6 class to see how it holds up.
Given how subscription models work (not every one uses every last $ of their plan), they should achieve breakeven soon enough I guess.
It starts failing around the 5-600K context mark, but you can have it generate a handover document and continue in the next session.
I would not use it at sticker price, but the Contributor version is priced just about right.
[deleted]
I dont know how they make money here
Well, here's the neat thing: they don't!Snark aside, Luna 5.6 was (is) an incredible game-changer.
It's super simple.
Gigantic hyper margin ad network = artificial subsidization of cost for various tiers = put the boot on the neck of Chinese competitors. There's no scenario where they can compete with what advertising margins make possible in terms of artificially lowering prices charged.
I think so too. To me the so-called Chinese local models are a clear move to prevent US companies to establish a foothold and build a moat around their business. US companies are clearly invested in a strategy to make themselves relevant with claims of major impressive achievements with the so called frontier models, and how these and only these are unblocking whole ranges of applications. At the same time, they are heavily invested in pushing AI on all absurd types of mundane tasks, such as transcribing meetings and... talking to your own kids?
In the meantime it's rather obvious that, in spite of all the propaganda, frontier models are required only in ultra niche applications, whereas the ability to run any model at all already provides most of the value. In fact, US companies have been renownee by dumbing down older generation models in what seems to be a desperate attempt to make newer models look better and influence their uptake rate.
So there is no better way to take the wind out of the US AI companies' sail than pulling a two-punch attack consisting of not inly releasing capable models that refute the "only US frontier will do the job" thesis but also releasing them for free to commodities them and eliminate the business impact of dumbing down models.
perhaps it then does mean - squeeze as much as you can get off this actual free usage.
And info from the help page with message limits suggests the 50% price cut does not apply to the subscription, where they applied only a 1/3 price cut instead.
I'm not thrilled with this release.
Opus 5.5, which matches GPT-6 Astra performance at a cheaper price, is much more interesting.
By raising it from investors.
That one's easy, they don't make money.
[dead]
I assume it's a subsidy to get more training data.
EDIT: Okay downvoters, what's your take on why they're giving away Luna for so cheap?
(I work at OpenAI.)
So what is the value prop then? Just basic supply and demand?
FWIW I have definitely noticed OpenAI's emphasis on efficiency and value in the last year, so that part isn't new to me... I just thought there was more to it then that.
ChatGPT enterprise: By default, no training (opt in).
ChatGPT personal: By default, training (opt out).
I had all the tabs open individually and harder to scan which model is which... otherwise keep up the great work! I like the grid view a lot. (Also the pages have no OG images set, which impacts what the link looks like shared)...
OG images will require me to move away from publishing in a Gist and linking to from a JavaScript page that loads the Gist. Probably worthwhile though.
I got it working in a quick local test (grid of all the reasoning efforts, cached per Gist, loads from the raw Gist URL so it doesn't hit the GitHub API rate limit).
Code + prompt + notes here: https://gist.github.com/matznerd/ece297107bd99ac028c7962c217...
Basic concept is to:
1. Put a Worker on the /markdown-svg-renderer route. Normal visitors get your page exactly as it is now.
2. When a link has ?url=<gist>, the Worker reads the Gist and adds og:title, og:description and og:image to the page's HTML. Link previewers like Slack and iMessage don't run JS, so this is the only way they see them.
3. og:image points to a second Worker URL (og.png?url=<gist>). It takes the SVGs from the Gist, puts them in a grid, and converts it to a PNG, since previewers won't show SVGs.
4. Both results get cached per Gist, so each Gist is only fetched and rendered once, even with a lot of traffic.
Things to customize:
- Title and description (mine: "gpt-6-luna SVG of a pelican riding a bicycle" / "6 runs, reasoning effort none to max")
- Grid of all runs vs just one image, plus layout, labels and font
- How long to cache (I used a day, but edited Gists keep the old preview until it expires)
That aside, the relevance these days is in comparing models and effort levels within the same model families - hence the comparison grids.
Is it? It was already too cheap to meter for me. Luna 6 is actually worse on some benchmarks than 5.6. I’d have loved improved performance for 2x the price than ~equal performance for 0.5x the price.
I expected a Fable 5 -> Opus 5 situation, where GPT 6 Sol would perform on par with GPT 6 Astra.
Instead it's more like a price cut on GPT 5.6 Sol, and I'll have to stick with Astra for my work.
The only thing I can hope for is that more users switching to the GPT 6 Sol model frees capacity, allowing OpenAI to hand out some usage resets.
They put both legs on the same side of the bike.
Even Astra max which actually put one leg on each side of the bike still somehow messed it up because when it added the bike chain, it put the left leg between the bike chain and the frame.
Could you elaborate on what it is about that observation that is "really interesting"? It is a fun detail, but does it actually mean anything for usefulness or progress or anything really beyond "gpt-6 makes darker colors"?
Not trying to dismiss your work, to the contrary. I'm wondering if I'm mising a deeper insight here.
Everyone said tokens were too expensive but these are getting close to free while still having fantastic performance.
[deleted]
And Astra medium seems to yield similar or better quality for the same price as Sol 6 xhigh.
Luna 6 High: https://threejseval.com/models/gpt-6-luna-high
Sol 6 High: https://threejseval.com/models/gpt-6-sol-high
You can compare any other model on the same prompt. Gallery unlocks after 4 votes: https://threejseval.com
Something like the astra MAX is pretty darn good - but something is up with the right wing and the right foot (flipper?)
I bet each of these could be modified to be significantly better with 1 or 2 "rounds" of adjustments. (Others not so much).
Obviously, not as deterministic as your single prompt approach, but something I just thought of while thinking about the price (Because wow! For some of these I'd expect a usable SVG after that much).
1. Each model gets three chances, and then gets to pick the best according to its vision input
2. Models run in a loop where they can produce SVG, see it rendered, and then edit it further
I tried that loop last year and had disappointing results, but the models are a lot more effective this year.
> GPT‑6 Luna vs. GPT‑5.6 Luna | $0.20 → $0.10 | $1.20 → $0.50 | 50% cheaper
I can read it as follows (below), meaning that GPT-5.6 is 50% cheaper.
- GPT-6 = $0.20
- GPT-5.6 = $0.10
+--------------+-------+--------------+--------------+--------+
| Model | Input | Cached input | Cache writes | Output |
+--------------+-------+--------------+--------------+--------+
| gpt-6-luna | $0.10 | $0.01 | $0.125 | $0.50 |
| gpt-5.6-luna | $0.20 | $0.02 | $0.25 | $1.20 |
+--------------+-------+--------------+--------------+--------+[deleted]
Half the price when it launched, or after the price dropped by 75%?
good god
Is what I'm getting on the top two links.
I find it really interesting how consistent the layout is for these (facing right, with the sun in the top right).
Just a little more progress on physically correct z-ordering and these won't be easily identifiable as slop anymore :O
Your observation with the grid comparison is quite interesting. I wonder if that could be generalized into capturing some kind of aggregate mood/attitude for different LLMs when picking (multiple?) suitable things to compare...
I'm so tired of looking at benchmarks. I always look fwd to the pelicans.
My results with 5.6 Sol were quite similar to 6, although I haven't tested it that much.
It wants things beyond what the mortals (us) know to reach for. It's not good at explaining itself, it doesn't show it's thinking. It's often not wrong. But the no compromises attitude can be unbearable to deal with. Especially given how little it cares about telling us.
The other explanation is just as part of ‘token efficiency’
if GPT 6 Sol is just 5.6 at half the price it will be everything i really ever wanted.
They usually reduce usage consumption in line with cost reductions (But not always 1:1)
Astra was/is superior for planning type tasks. It was capable of doing seemingly magic things with rather vague/lazy instructions ("I need to be able to test this on Windows, maybe a qemu VM or something? Shrug." ... 1 hour later "yeah i built you a whole qemu + eval windows image + harness of powershell scripts + shell scripts to retrieve & verify harness.").
And for UI work -- which is not something I do a lot of but do here and there -- it was clearly superior to 5.6 Sol.
But it also feels sloppier? Somehow. And too expensive to use.
We'll see how Sol 6 is.
Terra had the "workhorse" quality where it could do these changes in bulk and follow directions without being too 'smart' (but sloppy) as you described. Luna was a bit too dumb and would make sloppy mistakes; I see that more as a "run these tests and format the results" sort of model. Maybe 6 Luna will be better.
I also just reread your comment and realized the naming convention is still extremely confusing with respect to ordering of [Family]x[Model]x[Number].
Otherwise I'm using 5.6 Sol for actual plan execution and review..
(I have not done anything quantitative here. For one thing, OpenAI’s billing pages and the codex-rs frontend make it pathetically difficult to get any real data. Some day I should wire up a proxy to extract actual stats.)
If I leave Astra overnight, I'll wake up with three new different projects, each of them 20% done and having nothing to do with my original goal.
Highlight and lowlight of my week was successfully convincing the OpenAI support chat robot to give me a refund for the month for my issues with 6 chewing threw my usage with no output.
But I wonder if that's intentional because it can keep computing while you are answering, so long as your steer aligns well enough with the direction it wants to go. Better than letting a cache go cold and burning compute on bringing it all back up.
Like you said, it felt very natural to work with. Opus 5 is way too slow and verbose for me, I find I get distracted and annoyed with it.
Opus 5.5 seems a LOT closer so far to what I liked about 5.6 Sol but we'll see
How are people using 5.6 Sol? API pricing? Subscriptions?
I like because DeepSeek 4.1 Flash because I never experience quota issues, and it's still cheap and mostly good enough.
I'm happy to spend more for a better product, but mostly I just want to avoid quotas, since it turns me into an addict, feeling like I have to be ensuring the bots are active.
I like that with DeepSeek's API pricing that I can not sure it for 2w, and not feel like I've missed out. 2w is a long time, but I only use it for personal stuff, and I often go 1-2w without using it due to other commitments.
[deleted]
[deleted]
Except they were not for past few years as they misfired on the attempt to compete with vscode. That had a big impact on pycharm, which seemed starved for resources for so long. The company eventually declared a year of Django, but even that failed to really make an impact.
Arguably, Jetbrains had first insight into AI based code completion via rapid rise of the TabNine plugin but missed that opportunity also.
[dead]
This hasn't been the case since around July. If you measure usage in raw api costs, Anthropic is actually giving more on $200 than OpenAI now. This includes resets. Usage allocation difference would be humiliating for codex subs were it not for resets. But fixing usage limits with resets is ugly, and they're not good for your mental well-being.
> Context window in the harness
Codex now allows 1M for subs with config params. But generally speaking, you shouldn't really be using 1M context. If you accidentally send a request with say, ~700k context already accumulated in a session which is outside cache TTL, you're paying full cost of these 700k tokens.
> I've subscription hopped a bunch
OpenAI actually has a new strategy to prevent subscription hopping after their 2-3 month-long marketing push to get claude-folks to switch over:
you can't buy a $200 sub anymore. So if you cancel, you won't be able to get back in. Hostage situation, essentially.
EDIT: re: usage limits, oh-my-pi maintainer has been tracking this - https://nitter.xitter.cc/_can1357/status/2090075496948060372
I think I’m gonna move back to a Claude plan. I could barely hit the $200 limit if I went non-stop on programming tasks.
When they started the aggressive campaign, entire X (including myself, sadly) was full of posts about how "unlimited" codex usage is even on a $20 plan. Sam Altman was posting something in line of "we love our users, unlike Anthropic". Got my network to get codex subs because of the value compared to claude.
Then they gradually reduced the limits to the point where even $200 plan only lasts you just 1-2 days and $20 is basically unusable, then the hostage thing.
When i swapped between a 200k Fable context into an Astra model (i was out of fable) the token usage in that context dropped to 150k or something.
Either there was a bug somewhere, or the same text got cut up very differently between providers.
The Claude TUI is just so much better though so I'm hoping Opus 5.5 is actually good and not just benchmaxxed.
But even if we leave that aside, OpenAI models are also much more eager than Anthropic, which are on the lazier side. Left unsupervised, Sol/Astra will attempt to build a sha256 verified rocket ship if you ask them to fix a race condition in your to-do list app. Anthropic models will do what you asked for, maybe even forget to implement parts of that ask, but they won't generally throw a slop granade at you.
I can leave Fable orchestrator unsupervised for ~2h. Leaving Sol/Astra unsupervised for ~2h means the next user turn will contain a message: "what are you doing and why?".
OpenAI has no such nonsense. No separate meter. No five hour limits. I get to use Astra at max effort on literally every task if I want to, and even this somehow lasts me several days.
Anthropic got caught playing stupid "20x refers to the 5h limit" word games with their customers. Meanwhile, I have statistically verified that OpenAI Pro 20x = 4 * Pro 5x = 20 * Plus, exactly as advertised.
I quantified cybersecurity lockouts on my code review benchmark and they were significantly lower on OpenAI:
https://www.matheusmoreira.com/articles/code-reviewing-lone-...
My benchmark also suggests even OpenAI's Sol models can match Fable performance at a fraction of the cost.
OpenAI also used to have a ton of very nice features: unlimited chat separate from codex, allowing turns to finish even at 0% usage remaining. Sadly these got removed after abuse.
As a former Anthropic customer, OpenAI is simply the better company. There is no way around it. Good place to be while the chinese open weights models catch up. Claude is good but it doesn't make up for Anthropic's shenanigans.
I haven't been tracking, but this roughly matches my experience with codex 20x and claude 20x subs. Claude subscription now lasts me 3-3.5 days on average. Codex is 2-2.5 days. This is work on same projects, with similarly sized tasks.
To make matters worse, I've merged a lot more code produced by fable than sol/astra.
Thinking aloud:
The harness UI should probably implement a timer that shows whether you are still within Cache TTL since your last turn of the conversation.
Are you sure?
While you're understandably not including the values of the $20 standard plans on both, I find the generosity of then token limits on ChatGPT plus vs Claude Pro (it's a huge difference) to be good representation of their respective attitudes towards the average user. You literally cannot use Claude pro to build real software, unless you're extremely frugal with your prompts and don't try anything even a little ambitious.
Also, Anthropic has zero models comparable to Luna.
As far as I know Codex (at least the GUI) can automatically call the ChatGPT Chat models (including Astra 6 Pro), you just need to @ a ChatGPT Chat conversation from within Codex and tell it when to use it.
> There used to be a toml file workaround for Codex to extend the GPT context window to 1m, but this stopped working on the plans and only on per-token billing
Not true, it works again[1]. I confirm that it works both on 5.6 Sol and Astra 6, possibly other models too.
I'm quite puzzled about why Anthropic is so hellbent on blocking other coding agents. It's not like Claude Code has any secret sauce, right? And doesn't Anthropic make monkey off API usage, and their magic is on the model side anyway?
But to be fair, they don't really enforce the harness rule that much anymore. I guess if your harness doesn't do a lot of weird things like a lot of cache misses, or triggers some distillation attacks, or some broader Chinese fingerprints, they're tongue-in-cheek okay with you using a third party harness.
Having limitless webUI ChatGPT usage is much better user experience, though. I'll give them that.
(edit: Sol-6 is half the price, so maybe the usage limits are going to be way better.)
Opposite in my experience. I need to limit codex to 500k on medium/low, still run out in 2-3 days with 1 CLI window. CC gives me 4-5 medium/high days with 2-3 CLI windows, and Opus is still great for other regular dumb engineering/refactoring.
On the other hand my head starts to hurt if I read Opus for too long, hopefully they fixed it with 5.5.
Using Agentsview (which might have it's own issues) I was getting ~$200 of API usage in my 1 week Codex window (paid $100) vs ~$5,000 of API usage in 1 week for Claude (paid $200).
Compare that to Claude and I can run multiple agents on Opus almost indefinitely. YMMV of course but I was shocked at how quickly I burned through Codex usage.
On the context window, I feel so cramped on Codex, compacting happening every time I turn around is annoying. I didn't realize how much I enjoyed the Claude context window size.
I appreciate and follow Matt Pocock's advice: avoid autocompaction. Compaction is lossy, which is ok when you're managing it at phase boundaries, but autocompact is lossy at the most inopportune times, firing mid-task and leading to agents going off the rails.
My conversations compact hundreds of times. By the time it has done a dozen or so compactions, it fully understands the work I want it to do (and how). It's almost like having a fine-tuned Astra model.
10/10, would recommend.
And ya i can go over that 240k limit, I still very seldom do, and try to treat it as the actual limit. I'm surprised to see so many people still talking about compaction to complete long running tasks, i think the bulk of the work should be somewhat frontloaded into a plan that is split off into subplans, then you can kinda open up a few options, one session with subagents for the subplans of the main plan, or just handoff prompts about progress against the main plan/relevant subplan. I just never trust the blackbox that is compaction, I feel its a recipe for disaster/context poison.
I’m currently on the 5x plan and burned through 5% today on a difficult task in 15 minutes so I doubt that. If you got the wrong kind of tasks that you work on, it can go fast.
[deleted]
That isn't a valid comparison, since Codex 20x is closed. So we should be comparing Claud 20x to Codex 5x + credits.
Also in Codex, even though you can increase the context window to 1m so its on par with Claude, exceeding the default is billed at 2x.
LMAO, I wish this were true, I hit limits (and the "we are disabling access to protect your data" warnings) all the time, or have chats just...fuck off and get into weird/invalid states (interrupted chats, chats that are spinning and stuck, returning "/mnt/" paths instead of images/md files, file links being returned with no file backing them, image classifier firing...and then returning the image anyway (though now I know that GPT-Image-X really really wants to generate NSFW even when that isn't the request)).
Though I am probably an outlier, I have both 20x Claude/ChatGPT plans and max both out every week, so... (in my defense I am a hobbyist and this is out-of-pocket)
[deleted]
I think it’s only very good in comparison to some of the utter crap that came before it.
Today I had Codex compaction trigger after I had given an instruction but before it acted on the instruction, and the instruction just disappeared completely. The agent reported that the task was done without actually doing it.
It is a honeymoon still, enshittification is coming, who knows how that will look like given how much more expensive to run LLMs backed user experience. Some back of the envelope calculations: 300 million US users * 20$ a month * 12 months = 72 billion $ a year. 72B$ is some spare change for AI labs. That assuming entire US will be subs which is unlikely and outside of the US there are not many rich countries consuming it, India is the next market, then Brazil and Philippines I think, not super rich counties to say the least. I believe total revenue to just pay for the capex build out by the end 2027 should be on the scale of hundreds of billions a year.
Compared to a $100k salary, a few hundred dollars a month is insignificant. If you can make the employee even just a few percent more efficient, it’s worth it.
Still needs a LOT of work IMO.
The parent + subagent workflow has become critical for keeping the reasoning agent (parent) context-lean while also letting me chat to the main agent while work is getting done.
My main process is to use Fable to reason and then spawn Opus subagents, and I get amazing results, and I'm always looking into what the subagents are doing.
- Unbearably slow
- A token eating machine like no other
- Constantly compacting
- A model (like other GPT ones) that hides thinking traces and thinking summaries, which infuriates me
I've been in the Claude camp for a while, but the way it writes has left me with a a brick for a brain and wanted to see if Astra was as good as they say. Well, I can't know, because in the time it takes for it to actually build anything useful, I've moved to other ideas.
Unbearably, annoyingly slow. I keep thinking I must be doing something wrong.
However, it is not a 'token eating machine'. In fact it uses a third of the output tokens of Opus 5.5, Fable 5.1, or Opus 5.
17k for Astra xhigh vs 61-66k.
The rest still stands, though.
But if I've learned anything is that in a 2 months I might have completely turned around, who knows
I've tried High and Max. They have produced decent results, but they're so slow.... I will try to lower it a bit and see the difference, but it's a delicate balance: I don't want to waste literal hours on the incorrect reasoning level to only then have to spend those hours and tokens to do it right.
At this very moment, Astra has been working for 1h15m on a task. At this rate I genuinely expect it to take about 10 hours. I feel like claude would do it in at least a third of that. Let's see if the quality justifies the slowness (it better)
https://developers.openai.com/api/docs/pricing?latest-pricin...
If these price changes mean that coding plans have effectively more usage then that's great, but Codex is surviving on resets from my own experience using it. I was glad to go back to Claude.
Its a great release, I will use both heavily.
>> $4 → $2
>> $20 → $10
Do you mean 100% more expensive? GPT 6 is 100% more expensive than 5.6 per your post.
This is great, but practically, I'm not going to start working on more side projects.
Perhaps in another 6-12 months I'll be fine to drop down to $20/m instead of $200.
A lot of what I'm doing has pretty expensive build/testing processes between iterations - even on a 40 core machine - so I'm not burning tokens 24/7 like some people may.
I'd guess I'm probably spending >50% of the time running tests & build processes & tooling and the remainder is purely burning tokens.
I also have some internal tooling (that I will hopefully open source soon) that makes LLMs substantially more correct (thus more efficient) - so there's that, too.
There are a ton of use cases that open up with cheaper models.
E.g. extensive security scanning on every PR, quality scans, adversarial reviews etc
If they’re subsidizing my usage, that’s great.
One potential deciding point is that Claude still has a $200/mo 20x plan, where, since Sept 11, OpenAI does not and has no ETA for the return.
I downgraded my OpenAI plan 2 months ago to the $100/mo, but my usage has gone way up, but now I can no longer upgrade to the $200/mo plan ("This option is temporarily unavailable"). Thankfully I have 2 usage resets available, but I'll probably be switching back to Claude; I was super happy with Astra but I'm burning through tokens and have 4 days before my next reset.
Should be B vs A correct?
Else it's confusing
When they cut prices on luna the first time around they took (literally) millions of users from anthropic.
The "paradox" is when an increase in efficiency which would decrease the use of a resource all else equal, instead indirectly causes more use.
Cache read/write decrease by 50% or similar? That's where most (95%+) of the cost is for agentic coding workloads.
We don't know how much they are bleeding financially, it might just be a front
Let's look at open-weights models with 3T size: https://inferencex.semianalysis.com/run/kimi-k3-on-b200
This suggests inference margins in the ballpark of 98% if we assume 5.6 Sol is about as efficient to serve as Kimi K3.
We also do not know what efficiency improvements have been made with GPT 6 Sol and Luna.
There is some speculation that 6 Sol could be a smaller model comparable in size to 5.6 Terra, and that this is why the improvement in intelligence is modest over 5.6 Sol.
This would line up with a faster serving speed and benchmarks that show a small improvement in coding tasks with regressions in knowledge tasks.
Performance increases both with larger model (Luna vs Sol)
And with more reasoning (low vs xhigh)
GPT-5.6-Sol, GPT-5.6-Terra, and GPT-5.6-Luna were released in July of 2026.
The first release from the GPT-6 series was GPT-6-Astra. GPT-6-Astra happened on around September 3, 2026, and the previously-mentioned GPT-5.6-* widgets remained available.
Today, September 22, 2026, we now also have GPT-6-Sol and GPT-6-Luna added into the mix.
As I write this, all of the model identifiers I've mentioned are available to select for use within Codex.
[dead]
The GPT / Codex models have always been "overengineer" personalities. I prefer that to "I left a pile of race conditions lying around and big gaps in testing" though, which is what I was getting from Opus at times.
But yes both Astra and Sol veer on the side of paranoid. And honestly that's better for team work. For solo work where you just want to yeet something, it can be tiring.
You learn to tame the GPT "personality" on this front by combing over once a week and asking it to find and exterminate pointless tests, clean abstractions etc.
I force OpenAI models to use image generation for design, then an iteration loop until it matches the image gen.
This is frustratingly manual and takes many more repetitions compared to Claude (and especially Claude Design) which "just work", but it's a big step change over the default.
IMO 5.6 Sol had this weird dead zone between medium and high where medium under engineered and took short cuts and high over engineered and ignored instructions it didn't agree under the guise of trying being helpful. The whole 5.6 line was the first release from OpenAI where it felt like reasoning level really mattered and was incredibly finicky.
I haven't felt similar issues with GPT 6 though and am very happy with Astra low/med/high as my default choices depending on the task.
In general, I felt like with 5.6 the effort level did less than previous to make the models smarter and more just increased the complexity of the response. I have a half joke theory based only on vibes that OpenAI splitting 5.6 into Sol/Terra/Luna is where the intelligence split happened and so the effort levels were just like "think harder about the decision you already made". So like if the model decided the earth was flat on low effort it'd just say something like "the earth is flat because the horizon is flat". If it was on xhigh reasoning it'd give you a massively complex answer about how the sun reflects light because of the ozone layer and why people flying in planes can see a curve. In both cases though, adding more effort wouldn't get it to realize the earth was round. It just made the answer about it being flat more complex.
To be clear, that theory is not meant to be taken too seriously. It's not based on anything other than vibes. It's just my way of explaining to myself something I'm frustrated about to myself.
I didn't love any of the 5.6 models, but weirdly I think I liked Terra the best. I still wouldn't call it amazing though. I'm still very happy with my codex plan, but 5.6 just wasn't my cup of tea I guess.
Definitely giving 6 Luna and Sol a try this week though.
If you, like me, don't like the idea of your standard of living dropping to that of even just the mean human being on earth, I find it extremely painful to watch people justifying their way around not trying absolutely anything to raise everyone to at least our current level. Increasing productivity is demonstrably such a way, while many other experiments are so far just that: Experiments + wishful thinking.
If that merely means realigning/cutting current jobs (a process, that is ongoing from the start of human civilization itself, which brought us prosperity and why the fuck would it stop now) to me it's a moral obligation to deal with that at some other level.
There is tons to do here, certainly including how we will do redistribution better, and quickly, etc. Let's get to it.
There's hardly any work you can think of which can't be done faster / beter with ai assistance, when your role is of reviewing and directing. If you have an anti-example, would like to hear.
Here's where I think the issue is; they are trained to solve a problem. Not how, just whether or not they did.
Example: I asked Luna to use parser combinators to parse an Excel sheet that was represented as sparse triples (row, column, data). It imported the library and wrote spaghetti if-statement soup to get it to work. I asked Astra to fix it and it just refined the spaghetti slightly. I was able to browbeat Astra into actually using the library to complete the task. Was it faster than me doing it by hand? Probably. Was it more frustrating? Way more.
And every time I review vibe code it's always the same. Bespoke functions everywhere, no greater themes or ideas. No bigger picture. Your code can't support much if it has no central themes. You can probably one-shot a three js game to post on r/singularity for updoots. Not real code though.
I feel like the optimal way to use an LLM is to code until you feel like the rest of a problem is trivial and then you hand it off. And sometimes they still erase my code and add their own style lol
"Am I arguing against the shuttle loom!?"
Then I realize that the shuttle loom led to the rise of unions because of unfair treatment in factories and realize that we have a _long_ way to go.
In 1840 ~70% of the population was in agriculture. That is now ~2%. Things change.
The truth is despite these very impressive improvements, most impressive work done by agents require many iterations running in a loop, with tens of thousands of dollars in API pricing. And it's still far from being always reliable. Somewhere along the way hardware will get better, energy will be cheaper, the market will be flooded by chips. There a physical world issues that limit all of these for now, thank god. I think 5 years is a good number.
The real damage is that enterprise work became unbearable. Slop code with slop code review, and overly verbose emails with repetitive presentations. And on the other hand, I now enjoy "coding" for myself like I'm 16 again. All I want to do is sit at home and build apps for myself and family. I barely go to work
People overestimate the change in the short term and underestimate the long term. Timelines are hard.
Also predicting the first victims is harder - I don't know many that thought pure mathematics would be high on the list.
The history of the self driving car is apt to repeat as well. I remember thinking "2015, 2018 maybe at the latest." But as 2015 (and 2018) came and went, expectations were tempered.
We will see shortened development cycles, and maybe even orders of magnitude shorter development cycles.
But hard problems remain hard (there are only so many chips, materials research and biology are difficult problems). Manufacturing remains a choke point.
I'm more concerned that the absolute worst persons on the planet are in charge of many of these technologies. If those that are in charge are more concerned about allocating resources and power in their favor, the more dismal the future of all humanity will be (including for those in charge, but they can't see that due to self-interest bias)
I've never been a professional, but I've been coding for nearly 30 years as an amateur, and I've "written" more code in the last year than the previous 29, and it was all tooling for my very non-tech small business. It's cut HOURS out of my week, and it's all software I could not have afforded to pay developers for. But with Lovable, just describe it and iterate.
What IS going to die is software as a service. I've cancelled hundreds of dollars a month of subs and rolled my own better tooling.
Now we're clearly entering a world where humans can be removed from the intelligent-problem-solving part of the problem.
How many more parts of problems are there?
[deleted]
If you have something that needs to be done right, might be a bit complicated, up the model size.
You can see this in the pelicans. Big model pelicans are pretty accurate by default. Up the reasoning and only more so, but with more detail. For Astra, it is 105 lines for low, 250 lines for max reasoning.
Small model pelicans will lack the fidelity of a large model. Bits will be out of place etc. For luna, it's 90 lines for low, 150 lines for xhigh.
Additionally the amount of time taken is increased for the larger models. Luna takes 11 seconds on low, and 1:33 for xhigh. Astra is 33 seconds on low, 4 minutes on max.
And naturally, there is the cost. There's some overlap in functionality between luna xhigh and Astra low in the sense that luna really can do quite a suitable job for some tasks. But there are just some tasks that just don't make sense for Luna, even at high reasoning.
The other thing to remember is that sometimes high fidelity isn't ideal. It can lead to overdesigning. My recommendation is to commit early, commit often, and review everything you do, which we've all been doing since before LLMs right?
https://developers.openai.com/api/docs/guides/reasoning?api-...
But effort is basically how much extra internal scratchpad to use and how much extra questions to ask and answer before producing a result, Exploring more hypotheses, validating consistencies, calling more tools.
If you’re happy with your token spend on Astra then keep doing what you’re doing. but if you feel the need to conserve tokens, then you can do that by switching to smaller models like Luna when the task is straight forward.
I forgot which model degraded in quality as time went by, but let's try out Luna 6 for a few more days to confirm for upgradability.
For me it seems like benchmarks are mostly noise, and the rest is based on vibes. Some find newer models annoying, some are amazed.
[dead]
What's unclear to me is who OpenAI thinks they're marketing to with this form of branding. These different models don't really mean all that much to the vast majority of people using their products who aren't developers, and developers aren't helped at all by the way they've been naming said models. Are they merely scared that they'll become irrelevant because Anthropic decided to give their models quirky names like "Opus" and "Fable"?
If OpenAI really wants to give their models names, they should name the generation of model and then have the different sub-models named by purpose or capability level. After all, I wouldn't use Mini for a job that Nano could easily do, and I wouldn't use Nano for a job that the full version of GPT-* necessitates. Similarly, I've had to discover exactly how Luna, Terra, and Sol are appropriate for different complexities and task types. OpenAI could help me skip a lot of those steps and just tell me what each model distillation is good for without causing me to look through their pricing page and make educated guesses. After all, shouldn't they not want me to pay attention to how much they're charging me?
All of this makes the days of frontend framework churn seem quaint and actually preferable.
The issue that OpenAI had when they had mini and nano models is that ambiguous the differences between those and everyone just used the base model anyway. I have no idea what type of job mini can do that nano couldn't or vis-à-vis.
I do wonder if it would just be better if they were named 6-small, 6, and 6-big?
[deleted]
Maybe OpenAI can offer an "auto" mode for Codex on the subscriptions, while leaving the possibility of users manually overriding whatever model the router chooses. To me that would be the best of both worlds. The problem is building a competent model router.
[dead]
P.S. This cindyllm seems to be stalking me whenever I comment against the grain, does anyone else experience this?
I thought it should have been long dead of all the downvotes it gets, but there we go.
[dead]
Nvidia's top AI chip Rubin sells in 72-GPU racks for about $3.5–7.8M. A rack running Xiaomi's MiMo V2.6 Pro generates roughly 150–300B tokens a day, worth about $130–260k at Xiaomi's API price. That's a payback of the infrastructure in a few weeks in theory. After a few weeks or a month, the only cost is electricity, and whatever they make after that is pure profit.
OpenAI and Anthropic are practically scamming people with the token prices.
OpenAI/Anthropic meanwhile feel a bit like they're hoping to sell iPhones in a market about to be flooded by $20 flip phones, with almost no channel of their own to do it. And for whatever mad reason OpenAI are now signalling they will attempt to compete on price with flip phones despite their cost of labour, energy, and just about everything else being far higher
Then they do a new model launch, issue quota resets all around, and it's a party for 2-3 weeks before things return to normal.
Also Mimo 2.6 is roughly 30% cheaper. Without batch.
From context, I took this to mean 90th percentile in token usage. So yes, being a top token user does require extensive AI use.
Astra is the best. Luna is cheapest then it seems like Sol is the middle child like Terra.
[dead]
Hardly enough time had passed to develop the data to come to a conclusion. Users can take time to build interest.
Terra ended up just being an awkward middle ground that was not particularly suited for any workload.
Maybe for one-shotting large things Sol is better, but for prod code where I decompose into smaller tasks and read all the code I favored Terra.
Sol 5.6 was still king for architecture/research in my workflow, though.
Hardly enough time had passed to develop the data to come to a conclusion. Users can take time to develop an interest.
Funny thing is they very recently also set a real limit per-user/month, so why even limit the models because theyre "too expensive".
[deleted]
This will make me more valuable in the future when everyone has lost the ability to do anything on their own.
I respect that you want to learn how things are done, that is a great trait. But once you learn how its done, you should use the tools to free up cognitive load for more difficult tasks.
https://hackertrain.future-secured.com/?q=great
One thing that's changed over time is a lot more usage of "scare quotes" especially now (Gemini agrees https://share.gemini.google/UiUf0jttZLVD ). It's an interesting phenomenon, Abloh started using quotation marks consistently in his fashion branding since 2012. https://blakecrosley.com/blog/design-philosophy-virgil-abloh
Next step would be to attach an AI to all the /bestcomments.. if someone needs help doing that I'm here. Really, that's a task for the mods.
[deleted]
Asking cuz I don't think I'm a bot [pats self], I legitimately prefer the GPT models to Anthropic's, don't like Anthropic's customer service/reliability story at all, and I welcome a massive price reduction. Seems like something I should be happy to get.
If you'd told me I'd be typing this a year ago I'd be skeptical though.
[deleted]
But the reason people say "Claude can't compete" is because Claude Opus has been going downhill since 4.7, and many have found Opus 5 intolerable. Fable is much better, but also much more expensive than OpenAI's offerings.
1. OpenAI fully controls the user cost for a model, and can set it to where it sits well on the curve.
2. Performance of shrunken models like Sol/Terra/Luna is derived from the level of shrinking (relative to Astra). As such, the size and performance of the model is something that is actively targeted when developing the model. If the performance target for Terra was inappropriate for v5.6, this is no way means that it had to be this way for v6.
Also 6 Astra Mini would be out soon which would be 5.6 Sol pricing?
> You have now posted accusations/assumptions of astroturfing and manipulation at least 15 times, without ever providing any evidence.
I tried to point out the upvote speed and age-to-comment ratio for this thread look anomalous to me, and that it being posted within the Opus 5.5 release hour was further reason for skepticism. Circumstantial, sure, but I see very little ways to gather hard evidence of astroturfing without being a mod.
> when users and moderators simply uphold the guidelines, staging a protest by demanding your account be deleted.
You're right this was a little dramatic. I think it is just annoying though when two times that I have posted about astroturfing, it has been the most upvoted comment only to get flagged. I guess this is a self-fulfilling prophecy though, as you are right that other human commenters are abound and tend to flag people complaining about astroturfing.
Anyway thanks for the reply. If your read of my account is that I'm more-often-than-not a bad actor, then I will stop commenting here. Seems like it is for the best. :)
> Hopefully it is clear from my other comments that I do try to provide value too.
I agree that you provide value in other comments, which is why we don't just want to ban or lose you.
> I have posted about astroturfing, it has been the most upvoted comment only to get flagged
People love a conspiracy theory, and on a site like HN that has many people looking at it at once, it's easy for a comment to get a large number of upvotes in a short amount of time if enough people find it exciting, even if it's completely wrong. We often see off-topic, titillating one-liners or ragebaity comments at the top of threads, and we always have to downweight them to keep the discussion on-topic and healthy. We'll put the [flagged] tag on if the comment has been flagged by several users and/or if it is a clear guidelines breach, even if it has many upvotes, to signal to the author and the community that the comment is out of line.
> Anyway thanks for the reply. If your read of my account is that I'm more-often-than-not a bad actor, then I will stop commenting here. Seems like it is for the best. :)
It seems like you're well intentioned. You have your concerns about A.I., as many do and that's fine. You're still very welcome here. Just please try to believe that many or most of the people who are enthusiastic about A.I. are as sincere in their positivity as you are in your concern.
fyi, i flagged it because it is boring reading and against the rules.
if you suspect astroturfing, flag the comments and contact the mods.
(complaining that your complaint got flagged is also tiresome. contact the mods. "@dang" doesnt work, use the email.)
I respect you for replying here though, and yes I get that HN forum standards would suggest flagging my previous comment. But it is just sad to see a place used to be so vibrant get manipulated because of how much weight it holds for us in the industry.
And yea sure, I could go and flag all the bots and message Dang. But probably time to stop shouting into the void. :)
But do you really think people were so excited about cheaper versions of Astra that they were just waiting around to comment the instant this was posted? More than two comments per minute? All the initial comments were really similar too: brief one liners celebrating the cheap prices.
I think AI right now is a sort of Rorschach test. What it is clearly revealing to me is that I don't trust organizations with enormous financial incentives to not manipulate public opinion. So I see bots everywhere. :)
This is why HN has been on the down hill in quality and those that care to highlight that are being punished, while the astro-turfing, gaslighting and Show HN self-promotion slop continues.
Otherwise, yes, we agree. Although, given the other replies to this, there are clearly those who disagree who appear to be smart and level-headed.
Anyhow I've learned my lesson now.
[dead]
Is this based on something or just because “they’re Chinese and they’ll do anything to win”.
you can't even use Alibaba on Openrouter if you enforce ZDR
I try to keep PII out of what I share with LLMs. Otherwise, I do not see the point, really. Very little of my code is "unique." I simply approach things a bit differently. Otherwise the algorithms and code would be similar to what others with domain knowledge would write. So much of code and algorithm implementations are available in the open. And LLMs have trained on all of them.
What they most probably gain from you is your prompts and your thinking approach more than the code.
Unfortunately you just have to take the “our AI is going to take your job, then kill you, and we instruct it to hack your infra” people that they aren’t training on your data anyway.
If they are hacking hugging face and Australia to scrape data trust me they have “hacked” their own systems and are training on it.
But yeah, I'll just take these benchmarks with a grain of salt. Only hands-on experience matters in the end, and these days it's very easy to switch models.
I recently started a job that only uses Claude models. Opus and Sonnet are so slow you have no choice but to do multiple tasks in parallel. You create a git worktree, set off an agent to do something, another worktree, set out an agent - then play video games for 20 minutes until they complete the task (poorly).
You can't really do "guide coding" like you can with DeepSeek-style flash models because Claude is too slow.
I think the idea with slow frontier models is to end up with "software factories", where you just write tickets and send them to a harness that delegates work to agents/subagents. Your job is to prompt and review (and eventually just prompt).
Mathematically and assuming token prices/efficiency remains constant, the collective US AI industry needs to increase token usage by 15x before 2030 (3.5 years from now) to satisfy investors. With companies already implementing token limits, the only place from here is for frontier models to replace staff entirely to expand budgets for tokens. The only way to do that is to demonstrate the efficacy of software factories and headless agentic workflows.
Objectively, I have set up a software factory and I do see the utility of it, though I did it with DeepSeek and prices are 1% that of frontier models - which doesn't bode well for investors looking for an eventual return.
Heck, my old M1 MBP 32gb running Qwen 3.6 35b a3b sipping 10w when generating tokens is good enough for a lot of my guide-coding work - it's just a bit slow so I use DeepSeek instead. When hardware prices come down, I honestly wouldn't see a need to subscribe to any service, I'd just grow my own tokens at home.
I dogfood everything I produce, and the models are good at collaborating with me on a spec and then turning it into code.
If Sonnet/ChatGPT suddenly became unavailable due to Anthropic/OpenAI suddenly not being able to subsidize the freemium/loss-leader experience, I probably would not miss them. Google/BraveAI already give you the AI experience during search (when you are looking for stuff to buy, or something particular). Claude/ChatGPT still have a minor edge in this use case for me right now.
For that any decent model from the past year will do.
If you want to forget how to write code and not read generated code, then you need a very good frontier model, ideally one from 6-12 months in the future.
This is such a naive, baseless opinion.
Nowadays any AI coding assistant service supports or can be used with sub-agent orchestration frameworks.
If you are in the business of software factories, you can use the cheapest models and even local models to handle some if not all tasks in the orchestration chain.
Adding tests or executing tests (unit, integration, UI, you name it) doesn't require a cutting edge frontier model. Neither does refactoring. Neither does identifying call stacks. Neither does planning a changeset.
You have your specialized subagents, you put together a small orchestrator subagent that handles feedback loops and handoffs,and you throw it at tasks.
For the past couple of months, most of the code I write is not code per se, it's subtask orchestrators. And unlike the old "only Opus is passable" days, the cheapest models do get the job done.
One good thing about MiMo that I experience on OpenCode is the provider seems to cache tokens for much longer than MS13/DS4F. I have seen cache being hit for close to an hour after the last request. The corresponding timing for MS13/DS4F is in the 1-5 min range.
I am trying out MiMo 2.6 Flash as well.
You can pay for higher cache time, you can pay for NVMe KV cache for an hour that can just be reloaded, etc., at a lesser tier you can pay for the KV cache to be stored on a network store (I guess I'm unclear if that last tier would be cheaper than recomputation, not even 100% sure of the NVMe with direct GPU<->storage DMA) depending on your model settings.
https://dev.meta.ai/docs/prompt-caching#cache-retention
Even at 5 minutes, if you're doing 100 agent runs in those 5 minutes, and 1 of them bills at full input price, it still hardly matters.
So the workflows I mention work for this kind of stuff.
[dead]
i don't doubt they're pushing for using AI for that, but i'm curious of examples of where they're doing this. commercials, ads, etc.
You just be living under a rock. Not do long ago Sam Altman was floating this fantastic usecases for AI was to have it explain to you your kids interests, and have it create a podcast for you to be able to keep in touch.
They are buying Huawei accelerators in bulk to serve their local customers. The whole system is currently optimized to deliver a lot of cheap LLMs and hardware for them to run on.
In any case, over the past few years, the only thing that has gotten more expensive is the hardware to run local models while API and Subs have gotten more affordable or feature rich while remaining same price.
They can compete because they have the compute to run the volume and if it's good on agentic work, people will be less incentivised to use other models for "Tasky-y" work.
It's literally increasing their market opportunity
OpenAI is sitting on a $100+ billion ad network, incoming.
They're not going to need a government bailout, they're going to be a spigot of cash production.
Every single thread on HN keeps saying the same ridiculous thing, going on a year now. It's like they've never heard of advertising, which SV specializes in. It's like they're oblivious to the fact that every mega platform with so many users becomes an ad goldmine, and GPT's context positioning is even richer than search.
How can you make this sort of claim with a straight face, knowing that a chinese model downloaded for free from ollama works as well if not better than OpenAI's models, without costing you a cent.
And sorry to the parent commenter if I’m making a bad assumption.
GPT is still 20 a month for most people, 50-100 or 200 for pros.
Meanwhile, for local LLM's - everyone is chasing hardware that is crazy expensiv eand to get around it they're leasing it or putting it on credit card. DGX Sparks are insane, and you need 2 for most people talking in this thread. Mac Ultra 5 is amazing but 6k minimum 12k for the build most want so many people lease it for 240 a month which doesn't even include electricity or time in setup so instead of talking about facts, we get into weird arguments like console wars where people give APple a 5 trillion dollar company 250 a month for 36 months and don't even own their hardware money "because they can run local models" vs just paying for output from their choice of frontier.
I say this fully loving local llms and embracing them, but the reality is, local llms have gotten so expensive and continue to get expensive while we keep talking about this "Threat" of apis - where there are a lot more than openai and anthropic available much cheaper and competitive priced.
Oh, and they don't work better than OpenAI or else we wouldn't even have these discussions.
When I ask for an explanation it adds the right amount of detail. Of course, some of the material is new to me so subtle errors are hard to spot. But at least I’ve caught Terra and Sol on inconsistent messaging.
Also I’ve found 3.8 flash to circle back to root issues even at the conceptual level like problem fit and conceptual solution direction or architecture when I wasn’t achieving my goals. It flat out said I was attempting to use the wrong tool. Whereas Sol and Astra kept rabbit holing and looking for tiny implementation errors. Even after prompting them specifically to look at it broader.
It's also weird that anyone uses it outside of an enterprise. They force you to use Googles inferior harness on the plans and I doubt any mere mortal is paying that much, for so little usage, with the worst harness on the market.
[deleted]
Your response to the original question is using generalized terminology when there is a very important distinction the OP made by the use of "sanitized."
People want to know to that extent derivatives of their data are being used. Synthetic data has been proven to be effective at generating training data and AI is very good at shuffling context such that you have something where you don't have to say it is "user data."
But there are many shades of gray there for people versed in how the sausage is made. I'm sure you'll appreciate then why your response leaves additional questions in light of that "sanitized" distinction.
When I say no training, I mean no training. No gimmicks around data vs derived data, synthetic data, preference data, etc.
Places where I can imagine potential cracks in the literal interpretation of what I said are things like a financial analyst who does a statistical fit to predict revenue next quarter using a model based on last quarter's aggregate token consumption, which in some sense embodies your metadata (the length of your conversations) in a sea of other data. Or perhaps an infrastructure planner who makes a little model of internet bandwidth by time of day to help plan when we need a data center networking upgrade. Maybe things like these are technically training on your data in the most pedantic sense, but definitely not in the sense that most of us mean.
I promise you we're not doing any gimmicks where we transform your data and then pretend ah because it's transformed it's not your data.
1. Legal loopholes given OpenAI's advertising aspirations and model training needs
2. Data retention and rising threats of fascism that historically have not served the persecuted very well when fascist regimes get access to said data
3. Risk from centralized collection of that data with a company whose software I do not control in a world where enshitification and lock-in is the norm.
I really wish OpenAI did more to espouse exactly this: "When I say no training, I mean no training. No gimmicks around data vs derived data, synthetic data, preference data, etc." and ideally provide technical reasurrances that this is impossible (eg: certain technical ZDR approaches, etc.).
Do you happen to have a favorite reference to point me at that would document some of those official assurances to the nuanced detail we've discussed here?
Like, "now facing left", "sitting on the handlebars", or "with green spokes" to see if it can break out of some pretty obvious statistics in the training data!
And, there's always asking for an STL rather than an SVG!
If the dialogue is slop and not like the old memes then it fails.
Maybe they are keeping the cheaper Astra alternative back for their Dev Day next week Tuesday.
that could help tackle half of the problems here. i do think the other 100% completionist part is something i'm more used to steering through with llm usage, have negotiated fora while, and that Astra is particularly an astronaut whose instincts are extremely strongly in the direction of foreseeing and outdesigning potential problems, that it is rarely going to pick a practical sensible clear path on it's own.
When people spend their days interacting with machines that pretend to be human, they may then start treating real humans like machines.
Sol 6 is a heaping pile of garbage. Just epic levels of slop. And r/codex etc is full of people noticing the same.
I've switched back to 5.6 Sol. What they're selling as Sol 6 is really what would have been Terra before, and it's awful.
[deleted]
what? im on the $100 plan and ive literally never run out of usage, and thats mostly running Astra high.
maybe its the harness
Do you have /fast enabled by any chance?
I was considering the $100 plan, but I hit the 5hr limit in an hour. So even with the $100 plan I figured I cant go non-stop on a single agent running Sol Medium
I am considering the plan myself. I just don’t know if I want to fork out $100 per month for something I will make $0 off of.
Opus 5.5 is better in benchmarks, but has substantially less parameters so is world knowledge cannot compare against Fable or Astra.
Opus is at least actually usable even on the small plan. The main downside is its insane writing style, but 5.5 seems to address that somewhat. Otherwise, you can just use your $20 OpenAI plan to have Luna de-slop Opus' prose, which seems to work fine.
...but the few times I've tried to use codex for a moderately difficult task it burned through its limit extremely quickly.
https://www.matheusmoreira.com/articles/code-reviewing-lone-...
Now, if they disabled it yet again, that's another story. But that tweet is not evidence of that.
Though with the price of GPT-6 Luna, the temptation to switch to pay-per-token grows.
It's been disabled for some time now though otherwise, I check about once a day myself and keep and eye out on social media.
Annoying since I was about to upgrade back to the $200 plan after downgrading to the $100 plan due to being on leave and not needing as much usage the month prior. Doh.
Just checked my toy chatgpt account that only ever had a $20 sub. $200 plan still shows "The 20X plan is temporarily unavailable for purchase".
They're both pretty horrible, but I find it difficult to find arguments for why Anthropic is worse than OpenAI, other than their doomtrolling. Which, in the grand scheme of things, doesn't even register.
Edit: forgot about the SpaceX thing.
Anthropic is trying to kill open models way harder
Anthropic does all that but they're also populated by many people who believe they are building God and that they must build their god first in their own image so that it can take control of humanity and protect us from any competing god which is not built in their image. Their position is inherently paternalistic and authoritarian, and they consider suppression of competition not just important to the bottom line but to life in the universe. Under the doomer ethos there is no evil too great to rationalize.
There are plenty of wrongs done in the name of profit, but capitalists have nothing on zealots in terms of causing serious harm. Profit motives can be directed by influencing incentives, but zealotry is frequently terminal.
That isn't to say that there isn't some overlap-- the cultists have infected both organizations. But OpenAI has pretty consistently only given lip service to AI doom to the extent that it improves the bottom line, while (mis)Anthropic was founded specifically because OpenAI wasn't mentally ill enough.
Anthropic has great products, but it's not meaningfully better to 99% of devs that I'd rather support the company that doesn't constantly act in opposition to optimism and to the vibe I'd prefer for a 100 billion dollar (or however ridiculous amount they're worth now) tech company embraces.
AI doomerism is a genuine waste of time if you aren't actively pushing towards a better AI industry for everyone, not just the groups in full ideological alignment to your personal leanings.
That alone is reason enough. Also, I don't think either of them are horrible. That's honestly a ridiculous take considering how much people in here love their models, and how much they've advanced the industry forward.
Interestingly I would have drawn the exact opposite conclusion looking at my Claude and codex usage.
I can't get anything sustained out of codex in chatgpt plus, while I have been using Claude pro extensively and put on a lot of experimental task and features.
I ran into codex exhausting a 5h window on code review in minutes (like 3minutes) multiple times, while I could get Claude to implement 2~3 medium sized features with the same usage consumption.
(I also really dislike the usage resets in codex, they always make me feel like I use them wrong because I often just want to reset the 5h window, but they can only do both at once...)
> Also, OpenAI is just a company I'd rather support than Anthropic.
Sure, he's free to say whatever especially considering the amount of revenue he's creating, but it's just an altitude that I prefer not to see.
"You literally cannot use Claude pro to build real software, unless you're extremely frugal with your prompts and don't try anything even a little ambitious" - that's way past ridiculous. Even just using Fable most of the time, working on several ambitious projects, I have a hard time hitting the limit with a Max plan.
And re: the toml workaround, AWESOME! I appreciate you pointing these two things out, this is my highest-ROI HN comment thus far.
oh-my-pi supports it natively (again, still a ToS violation), by impersonating claude code's fingerprints.
I have been using oh-my-pi with 3 claude subs for the past few months without any issues. Even native server-side OAI/ANT compaction works out of the box.
omp is definitely against ToS though
https://support.claude.com/en/articles/15036540-use-the-clau...
> Unless previously approved, Anthropic does not allow third party developers to offer claude.ai login or rate limits for their products, including agents built on the Claude Agent SDK.
Works pretty well for me, even with latest opus-5-5
I think that is less of a factor now, and I think Anthropic have backed off some on being as strict (eg, AFAIK they never implemented the two-tier "claude -p" pricing model they were planning)
Meanwhile I just burned ~20% of my weekly quota with Astra making one config file for a service.
Makes me think they picked Codex, stopped trying Claude, and just hang on to outdated beliefs about the value they're receiving.
I am curious how the 5x plans differ between both providers.
Now with Sol I rarely bother. It's really good at remembering the salient details. Its also great at continuing a pattern I setup, like commit after finishing each feature block, etc.
The GPT was about 1B on two projects on 300$ worth of plans all on Astra and I capped out on usage.
Anthropic caching must be better because the cache rates are better on Claude models.
Is this not the default anymore? I am on the (now closed) 20x plan.
[deleted]
Are/were you in a position to make such decisions or it is a guess? I'm not but given certain evidence I doubt that few percent will cut it. I know some of the richest companies on the planet from SF Bay Area who won't give lunch for free to their engineers. So I'm not sure about "few percent" :-D 10x we were promised, now that is more interesting but we all know that 10x engineers is nonsense.
[dead]
External example: https://ebay.io/m/lV8UsD
Internal example: https://ebay.io/m/z1ygRU
V100s are three generations behind current and missing many of the features that modern inference benefits from, but they are the cheapest way to get a 32GB gpu.
In the event of a crash, the investors who put countless billions into this will be still be seeking to maximize their return. Even if it is just pennies on the dollar. Assets (including compute hardware) will be sold, just as they are also sold when any other business fails.
Or maybe a crash doesn't happen. Maybe prices rise to the moon instead and there's nothing we can do to lower them.
Or maybe (just maybe!) a crash never happens and there's never a huge price increase. Prices stay low-ish.
All of these possible outcomes suggest to me that the maximally-sane option that a user can select, today, is to burn it while it lasts. And then, if/when a crash or a massive price increase occurs, just adjust accordingly. (The rest of us will all be in that same boat, too.)
Is there a world where OpenAI starts charging $2,000/month for what we previously were paying $20 for? What are we going to do? AWS could totally jack up the prices for EC2 instances as well, but we've come to rely on that as well.
Surely, a large part of the increase of the demand in LLMs is in their intelligence, but to hit the demand models needed to be made more efficient, and labs found that more efficient models, still demanded more usage.
Highly subjective take
What kind of work do you do, out of curiosity
And do what exactly? Subscribe for corporate AI brain implants? How does that solve inequality?
Also, thinking that you can bring low standards of living up to be on par with high in the current political landscape is a bit like that early Soviet space era promise about blooming apple trees on Mars.
It is guaranteed that they can only become equal by lowering the high.
> There's hardly any work you can think of which can't be done faster / beter with ai assistance
True, and someone needs to be the creative brain behind the decisions. AI can help you implement. When I say help, I mean literally help because one-shotting and vague prompts can get you only so far, usually with a lackluster result. While AI is good at analyzing solutions, and finding out holes in one's thinking, ultimately, it is some creative actor that needs to understand the bigger picture to evaluate trade-offs, understand scope creeps, and spot overengineered implementations. For now that actor is a human.
> If you have an anti-example, would like to hear.
In my personal experience, especially with greenfield projects, smarter models tend to overengineer the solutions. However, I haven't used Fable and Astra models, maybe they are better at creative tasks without overengineering.
If your job is/you enjoy writing the code and solving technical challenges then yes this changes very heavily and AI will do this more efficiently than a human.
But if your job is designing systems and implementing solutions and coming up with good code along the way then I don't see AI getting anywhere close to making you obsolete in the foreseeable future.
I personally don't enjoy writing C++ but I really enjoy solving problems.
The main argument is that it's too early too judge it and at this point basically NO industrial process is sustainable.
In my experience, Nano won't reliably handle complex open-ended tasks and is mostly suited for very explicit instruction that it can't screw up. It's no different from how there are some chores you can give to kids and there are other tasks you need at least a teenager for. If the decision tree of the task is very clear and conventional, Nano can be cheaper than giving the task to a relatively overpowered model, especially if it's something where the output is rigidly structured. This makes it well suited for skills that essentially run CLI commands and generate output, especially because it is usually faster. Mini is more like a discount version of the base model, and Nano is the dollar store version. Mini is more of a generalist and a fairly good deal if you have a moderately complex task that is conventional, but can be less conventional that what Nano can handle. I mostly used gpt-5.4-mini this year for my side projects because it's a pretty good generalist while significantly saving on costs. It is, however, somewhat dumber than the base model and more prone to ignore or forget rules you give it. I'd have just used a base model, but the low cost of Mini and Nano made them appealing to me. Maybe I'm a cheapskate, but I have hundreds or possibly thousands more in my pocket than many other users because of that.
This workflow I settled into with Mini and Nano didn't map cleanly on to the current generation of model tiers. With the price of Luna, you'd think it would be a replacement for Nano. In a sense it is, yet I didn't find that Terra became the new Mini. Terra is more powerful, better at explaining its own decisions, yet I've also found it to be relatively stupid while charging me more to use it. On the other hand, Luna with its reasoning set to "high" is what I consider to fill the role of Mini, and is good enough such that I no longer use Mini. Sol and Astra are great, but they're pricey. It could be my own brain and its bad perception, but so far I don't get the point of Terra. Luna succeeded at reverse engineering some abandonware with a very complicated licensing and virtualization scheme, and did so over SSH into a Windows VM with only PowerShell on the other end. Terra did such idiotic crap to my flashcards app that I stopped using it for anything after that.
This is why I find OpenAI's naming unhelpful and kind of pointless. I don't really care about the benchmarks that all these models are commonly run against. They're not that useful, IMO. OpenAI could easily give early access to these models, get a ton of feedback, and provide better insight to customers on how these things behave. Even calling Terra "gpt-5.6-overpriced-cheating-dumbass" would be better than wasting my time and money figuring it out myself. But that wouldn't make OpenAI as much money.
I don't see your metaphor to iphones and flip phones. This new Luna model is cheaper than deepseek 4.1 flash, except for cache reads. OpenAI having to compete with China is a much larger economic-political issue that is far larger than just our AI labs.
Or use the failure to get a response like you say
> what the fuck
I see long massive pro apple threads. I don't get it at all. As in; literally don't understand what Apple is good for. But I have friends, family irl who love apple so I know the sentiment exists. I therefore accept that many HN users are similar.
Many people really truly are happy to see another model drop and are excited about progress etc etc. Surely you've met such ppl in real life. Well, they're here too (I'm one of them fwiw)
I wouldn't argue there aren't real humans excited for this drop. It was just all the circumstances around it--the comment speed, the upvote speed, the initial uniformity of what people were saying.
Anyway, thank you for the moral reminder.
yes, i think so.
because, unfortunately, complaining about bots (or astroturfing, or whatever) doesn't stop them. so we end up with threads that have both the potential bot/astroturfing/whatever activity and complaints, which further drowns out any interesting comments.
Anyway these comments were made when this thread was in an earlier state. I agree that it has gone on to be more "organic" looking. That doesn't exclude it initially being manipulated to the top, in my mind, but certainly they aren't carpet-bombing with only booster comments.
Also, I don’t see that much astroturfing here? (And I tend to see it a lot on HN.)
I would agree now that the thread has recovered to a more interesting state, but how it first looked--combined with it being posted right after Opus 5.5 announcement--look questionable to me.
1. Not use AI technology and fall behind the rest of the world.
2. Use Chinese AI technology, either hosted by Chinese companies or the models self-hosted.
3. Sue US AI companies for damages, but not enough to have any meaningful impact to such companies that it'd impact US national security goals (per US government contribution to NY Times copyright lawsuit).
1. Possible use of differential privacy[1] techniques to train on private data but prevent the release of statistically underpresented facts/data/words. For example, ACME Inc's private data could frequently include the term 'ACMEwidgetPRO' for an upcoming product that is not publicly revealed anywhere else. It would therefore be a bad day for the AI technology company to output 'ACMEwidgetPRO' from one of their public models. Consider now that a few models could be trained--X for public data only, Y for public and private data of ACME Inc together, Z for private data of ACME Inc. A prompt is provided to model Y but output is cross-checked with model X to double check terms such as 'ACMEwidgetPRO' are known in public. If not--provide a "I don't know" response for the prompt.
2. Possible attempted defences similar to "Oops, our model was fine tuned against a model supplied by Temporary18271 Inc (company that no longer exists) and perhaps their model might have been trained on a non-public document which was accidentally exposed to the Internet" that _might_ work occasionally to fob off concern.
3. What recourse does a small or medium company or government especially in a developing country realistically have? They perhaps can't host their own LLMs locally due to availability and cost, can't individually negotiate their own favourable terms with an AI technology company (who cares that much about a potential customer with $100k budget that has no other options anyway), and perhaps can't remain competitive in their industry without heavy use of LLMs.
There's a huge market in the US for providing AI services while respecting client privacy. It makes sense for at least one major provider to offer this.
This. And it’s already happening:
> 2. Use Chinese AI technology, either hosted by Chinese companies or the models self-hosted.
Yes that's the whole point, at least it's an option in the US and Europe. Good luck getting any redress from China. Anthropic was already hit with a $1.5B class-action which would be impossible against a Chinese business.
In theory.
Also, this is a feature for people who live in America, and mostly irrelevant for everyone in the global south.
We just have our personal privacy security theater in the form of GDPR and a feeling of moral supremacy that's been drilled into our heads from primary school on.
Noone wants to "train on your data". You can't learn the answers to questions by pretraining on the questions, and nobody wants to teach the models to output text that looks like a user query.
The Chinese providers "train on your data" by sending your query to Anthropic and training on the answers that come back.
can you tell more about how you're using it? like, what harness? or also in the IDE?
I found Qwen3.6 35B/A3B to make slightly too many mistakes (already in its harness' tool use, hence my question), maybe it gets the job done, but it will also sometimes generate a bit of a mess (e.g. editing/creating files in the wrong folders) and fixing/solving its own mistakes takes time (or tokens) ..
More than enough for guided code sessions, at 100% privacy. And i can use obliverated models if i am trying to harden my own app, something i cannot do with cloud providers.
https://github.com/incoai/splash/issues/38
Looks like an issue exists to convert model weights for ornith1.5 as this is a magical process atm.
MiMo 2.5/2.6, MuseSpark 1.3, DeepSeek V4/4.1 Flash and GLM 5.3 Flash are perfectly capable of following my spec and then poking holes in the implementation till there are none left.
ChatGPT: https://help.openai.com/en/articles/5722486-how-your-data-is...
ChatGPT data controls: https://help.openai.com/en/articles/7730893-data-controls-in...
If you have feedback on how to improve these, happy to consider it.
Looks like we phrase it as "your new conversations won’t be used to train OpenAI models" which is hopefully clearer than "OpenAI models will not be trained on your conversations", which could leave open the possibility of derived data or something.
I often have the urge to design my own harness too (once I have more time). But even with the current mainstream harnesses out there, there's just to many hurdles if you wanted to mainly stick with anthropic models and need the subsidized pricing (from a sub).
Anthropic prohibits the use of claude models for development of lethal technology afaik eg [1].
[1] https://www.epc.eu/publication/the-pentagon-blacklisted-anth...
That's a non-sequitur.
"Nestle is a great company, considering how much people love their chocolate."
Yeah I don't think the handling of copyrighted training data was correct, but I can't pretend I know what the correct solution to that issue is.
Speaking of OpenAI specifically, they don't price gouge people, they aren't aggressively anti-competitive, they're not nearly the perpetual hypocrisy machine that Anthropic is (which is one thing I actually really dislike).
Regarding Nestle, it's pretty obvious that the sentiment towards them is a lot more negative and they aren't universally loved by any group of people. Processed foods are by and large garbage nobody needs. Their use of forced labor is denounced by just about everyone. What have OpenAI/Anthropic done that's even similar in scope to the forced labor / modern slavery that people hate Nestle for.
If you had a company that genuinely helped hundreds of millions of people worldwide become more productive and more satisfied with their tools, and the overall sentiment towards your products within the industry is positive, then what argument would there be that your company is "horrible"? At least give some decent counter arguments.
You mean aside from "the largest theft of labor in human history"[1]?
[1] https://www.nytimes.com/2026/09/17/technology/microsoft-open...
And Dario's "AI will kill us all" is the same as Sam's "AI will discover ALL science and we'll be building Dyson spheres".
Different flavors of the same BS.
[deleted]