Where Anthropic f'ed up was treating their monetization the way they treat model training. Turns out that success in experimentation is not transferrable.
They have tried to find the highest that the market pays for sota models; however, on the consumer side, this is just too confusing and unsettling:
"You can only use Fable for a week as a part of your plan"
"Be ready! You have to start paying per token!"
"Nevermind! we extended it for a couple more weeks"
"Wait, now it's up to half your usage"
"Ok, now its..."
Most people want to not care. We want our AI like electricity -- Kind of just there no matter how easy/hard is for the supply. You don't want your electricity company to be on the brink of cutting you off any second.
That's Anthropic. You don't feel they want to give you a dependable service for an, albeit premium, price. It's a constant bargaining game. That forces people to look beyond the walled garden. There, they find models that are fine... and without the shenanigans.
Yeah, I agree with this. The constant state of "...will the rug be pulled?!?" does discourage relying on it as a model and building a workflow on it. Anthropic used to just be a reliable thing you could play with. Now it's this constant source of anxiety.
It also didn't help that the government yanked it which adds another source of anxiety since OpenAI is on much better terms with the administration and the administration seems corrupt enough that they would mess with Anthropic if they got a big enough donation from OpenAI.
But anyway after Sol entered the picture, I don't think Anthropic can get away with this as much and I also think they're going to face a massive backlash from Max subscribers if they do end up ending the +50% promotion at the end of the month because Sol is a Fable peer and priced very competitively.
My wife's startup made the mistake of building her internal operations around Claude Team.
Then she hired a VA in the Philippines. Anthropic promptly banned her account without warning once the VA connected to the account. It took her weeks to get her account reinstated, at which point she had already moved on to OpenAI.
How many big tech companies let you talk to a human to get support.
Automation is wonderful to cut cost for them but for the users being unable to get support is a horrible experience.
But you cannot go elsewhere because they are the only player in town.
How can small companies with 1000x less money able to provide live support, but if you pay 20, 100, 200 dollars for a subscription you dont have a phone number to call ?
How do you know there wasn't account sharing though?
I've seen some amazingly dodgy stuff when hiring people from south east Asia, sharing a paid account with friends worth a months rent there seems milquetoast in comparison.
As soon as I started using Fable I was like, okay, this is probably as good a model as I will need for software engineering going forward. I still feel that way. I don’t need a better model, I need a faster Fable.
The thing I miss most about programming is flow, and the constant bouncing between terminal tabs sucks. I’d love to do one thing at a time, with Fable, quickly.
You may or may not like agents mode. I also hate flipping tabs, but I enjoy using agent mode with well named sessions. I still stick with a single session until I must move to another, then I leave them around for a few days until I’m sure I won’t need to pick up where I left off again.
I don’t think they were complaining about literally flipping tabs but rather just needing to context switch so often. This doesn’t sound like it helps with that.
There is GPT 5.6 Sol on Cerebras if you want to try that experience for an ungodly sum of money (not getting into GPT 5.6 Sol vs Fable, but only one is available on Cerebras) for an 11x speedup.
> They have tried to find the highest that the market pays for sota models; however, on the consumer side, this is just too confusing and unsettling:
The consumer side cheap monthly plans exist for the same reason companies like Cloudflare and Vercel have a free tier: When it’s cheap and easy to get developers familiar with the tools, they will push their companies to pay the real money for those tools.
It’s a hard balance with LLM serving because you can’t really make it free. $20/month is close to free, but the $200/month plans are in a difficult place where they’re big enough that many small companies pay for $200/month plans for their employees and ignore the enterprise features you get with the full expensive arrangements. So the companies are continually adjusting the $20-$200 plans to keep them from being reliable options for businesses, which is where the real money is.
There’s a short sighted cheering on of the 3rd tier and lower companies offering lower rates, but we’re already seeing them ratchet up the pricing and keep larger models closed after they get market attention.
> That forces people to look beyond the walled garden.
Every time I get "you used your quota, come back in 3 hours, or 2 days" -> that is experimentation time with their competition, leading to changed service plans. When they said "claude -p" will be billed at API pricing even for plan users I moved my harness off claude. After I integrated codex, then it was never going to be a full claude project again.
What business encourages users to try their competition and adapt their usage to the competing products?
I agree they’ve done a bit too many pricing / usage promotions and A/B tests.
The period was also marked with many billing bugs, like spending people’s usage credits for included Fable for a few hours (gave me a huge shock), but to their credit they refunded it.
They are fumbling the bag hard. AI's utility is for general purpose. The floor is rapidly improving from below them. With chatgpt I'm uploading all my day to day stuff. Meanwhile Claude is only for occasional super hard tech problems which are rapidly improving with being solvable easily by Sol. So what's the value add?
Their disrespect for their users is also another problem. You only get one shot to make a good impression.
If I got cut off from asking chatgpt what the status of a spreadsheet is for handling my personal bills due to my rate limit from coding work I would not be coding on it at all. Glad someone over there realized this obvious fact.
I'd go further and say chatgpt offering continuing service just with degraded model keeps me in their 'free web use' tier. I have API keys for coding stuff, but for most of my day to day use of public llm, I know I won't get 'blocked' using ChatGPT. Yes, the model may change, and I'll get 'worse' output, but usually I can't tell the difference. Using Claude for day to day stuff, I get completely locked out after so many hours. Again, I have API keys for Claude as well, and use it for 'pro' work, but day to day chat stuff... it's not my daily driver.
As Anthropic does this, OpenAI Is giving everybody resets like every other day now on Twitter.
I'm strongly considering biting the bullet and just ditching my $200/month Claude Code plan for the Codex one instead, especially because I keep running into my weekly limits (even sticking to Opus.)
I ditched Claude Code $200/month a couple of months ago in favor of Codex $200/month. The value is night and day.
1. No 5 hour usage limit
2. Weekly usage gets reset CONSTANTLY. It's crazy. The longest I've ever seen it go without a reset is maybe 5 days?
3. I don't feel like OpenAI is constantly trying to fuck with me. Unlike Anthropic. I would way rather have Sol all day every data, consistently, than a slightly better Fable for like, 1 prompt every 5 hours, and only when Anthropic decides to not treat me like a cyber criminal. Believe in yourself as much as Claude believes your CRUD app is going to hack the pentagon.
4. Getting access to image generation, though I don't use it too much, is a nice perk compared to Anthropic.
edit: Should mention that I had like 4 banked manual resets as well. It feels like OpenAI wants me to use their product, whereas Anthropic wants my money while giving me a nerfed experience
I got the AI ultra plan for gemini, the models are lower quality than claude and codex for sure but it comes with a pretty high limit for my purposes and I haven’t reached the weekly limit yet after about 3 weeks of usage. I have enterprise claude at work, the budget isn’t particularly great and I keep running out within a week tops with any serious work. Not even considering it for a personal plan with how quickly the tokens run out for even the mid-tier model/effort combinations.
I’m sticking with gemini for now given the 20TB cloud storage and youtube premium that comes with it but I’m still open to switching to codex (which I’ve had a long term plus plan for). If google keeps delaying the pro models for much longer or makes them excessively expensive, I’m likely to switch out.
> I’m sticking with gemini for now given the 20TB cloud storage and youtube premium that comes with it
I used to have the Google One with Gemini and Drive space and YT Premium separately for my family. A credit card expiration lapse lead to closing YT Premium and being locked away. Why? it was because Google One plan bundles YT Premium Lite as an extra. So they actively blocked me from getting Premium back for a month.
Now I moved to YT Premium on my wife's account and downgraded One to lowest tier. Never heard of a company forbidding users to upgrade their plans before this. I suspect Google really wants users on Google One no matter what they want. I was barely using Gemini anyway, already have claude and codex plans.
I've got the AI pro plan for Gemini and the limits as absolutely abysmal when using 3.7 flash in antigravity, If I use swarms then I can't even finish a single prompt without it reaching the limit.
I was going to say that the model of the electricity market OP is talking about hasn't existed in 15 years and it's only getting worse with more intermittent renewables entering the market. And no, batteries are not the answer because physics doesn't care how much greenwashing lobbyists do.
Yes, I feel like you can just sense the garbage coming. Age/identity verification, mandatory data sharing, morality policing, "Answer Engine Optimization" ads and influence peddling... it's going to be so painful to watch it all enshittify.
OP made the "electricity" analogy, and that's really all people want. I want to plug something into the wall and have it work. I don't want to have to worry that my electric company is going to rug-pull me because I plugged the wrong appliance in, or I didn't agree to some TOS, or I used the electricity to run grow lights for my pot farm, or this or that or the other.
The analogy is apt because I feel this is exactly what keeps Anthropic and OpenAI's owners up at night -- becoming the utility company the People want them to be. Ironically their behavior may accelerate their fears. And yes, the so-called safety features are ridiculously invasive and the worst is agreeing to have surveillance cameras installed in every room that occasionally detect any attempt to grow plants with LED strips as a pot farm. After a false alarm of almost having my ChatGPT account terminated for cybersecurity abuse and appeals auto-denied twice (I did nothing even close to hacking), I have started doing everything I can to decrease switching costs and thus the bargaining power of the suppliers and I'm doing the same for my company.
OpenAI is also now actively discounting their model. They just offered a free month to users. Both companies seem spooked by what I have to imagine is slowing user growth.
I don't think Anthropic wants a stable experience on their consumer subscription plans. It is just used for customer acquisition who will then ask their employer to pay for enterprise plan(assuming most employer care about data control) which is based on tokens.
Most coders don't pay for tokens themselves. It's just on reddit and HN you would think that everybody does.
I pay for my tokens for my own projects, at least when the ones Google seems willing to keep throwing at me for free don't cut it. I'd think that's not too uncommon, especially here where there's likely a high ratio of hobby coders (whether also professionals or otherwise).
I would guess that it's hard to compete with Google without a consumer component. Most enterprises already have contracts and policies with Google, so anything driven from the top-down is likely to prefer Google.
(And failing that, there was a real risk for OpenAI to be the default for enterprises.)
To defeat that, you need to frontline employees the chance to experience better tooling and models which is where the subsidized subscriptions come in.
Their truth is “we don't have compute and are working to improve capacity”
People would root for that
Instead they got people rushing to escape the permanent underclass until they have a mental health crisis just to beat the fake deadline. $100, $200, is a lot for those people
The fact is that demand for tokens at electric bill rates so far outstrips what can be supplied currently not just with frontier models, but with open weights cheap models too. Running an always on Deepseek flash agent would cost three figures a month at API prices.
Total costs sure, electricity only costs no. My two DGX Sparks run DS4 Flash at about 50tok/s concurrency=1 which is more than suitable; at about 150W total wall power when generating.
That’s about A$16 a month in electricity if I ran it 7x24x30.
>You don't feel they want to give you a dependable service for an, albeit premium, price. It's a constant bargaining game
We all know that the subscription prices are not at all sustainable for these providers. You all do, right?
Yes, they're struggling to segment the market and find a way to make money, and that basically relies upon emptying the pockets of whales. As someone enjoying a hilariously subsidized Max plan, I understand that, and I don't think they're trying to scam me in some way.
And both sides of this equation understand that the market is competitive, and maybe more competitive than they thought it would be. Like, would you rather they did pull Fable when they first said they would? Or that they'd cut quota? I wouldn't. But I'm glad that Kimi K3 and GPT 5.6 Sol and the latest GLM and Qwen and...I love that this has forced Anthropic to change plans. I'm not going to hold that against them.
> We all know that the subscription prices are not at all sustainable for these providers. You all do, right?
I don’t think it is fair to expect from people to know or understand this.
If you buy something or subscribe to a service there is a price tag on it.
You get X for Y amount of price.
That is how consumers conditioned for decades. They do not care what is your customer acquisition strategy. If Antrophic cannot provide reliable services on that price, customers will be unsatisfied.
DeepSeek is doing pretty much the same thing they said they weren’t going to change their prices after the 75% discount for the foreseeable future. That foreseeable future turned out to be two months.
They've put themselves in a corner. Fable was too good and they gave it away with the $20 plan. It had to be a big step from Opus 4.8 to show progress, and Opus 4.8 is GREAT at coding in many different domains.
But they're getting killed on token cost. They have to get people paying more for tokens. So then they put Fable in the $200 plan and release Opus 5. I'm suspicious of Opus 5. It is mostly worse than 4.8. It _seems_ like they nerfed it to create more distance between it and Fable.
So what have most of us done? Stayed on Opus 4.8. The statistics bear this out. 4.8 still dominates.
Now they're stuck. If they take 4.8 away, everyone will riot. If they make Opus 5.x better than 4.8, they disincentivize everyone from moving to Fable and most importantly, paying more.
Really, all they can do is take the L for now and just let 4.8 be the apex of the $20 pro plan for the foreseeable future while they work like hell to make Fable THAT much better that it earns the $200 to $infinity that they really want everyone to pay.
And they've hobble Fable and Opus so hard with their safety guardrails, I ask innocuous questions and tasks and they get flagged so often I gave up on it. I can get all the work I need done in GPT5.5 or 5.6 without the hassle.
> It tends to bail occasionally for me once auth-related code comes into play.
I have the same experience. I found that if I add to the prompt: “but it’s not security related, we only focus on the auth library design”, it “bypasses” the safeguard.
Yeah same. Multiple 20x max accounts, I typically hit > 50% of the fable usage limit on eace, literally never seen a guardrail. Maybe I’m boring? Maybe they have some kind of account reputation system?
I hit it once when I asked a question about whether butterflies remember anything from their time as caterpillars. I've never hit it for coding, but I also don't really do much related to security.
I kind of still chill out on Opus 4.6 too. 4.8 is good too. I go between them. Opus 4.8 is a little smarter some of the time. Their use of language is both very different from Opus 5. Some of the time I have a hard time believing Opus 5 is even related to Opus 4.6 and 4.8.
I'm going to be sad when they retire 4.6. It's not my daily driver but it's still my go-to when the other models are being stupid in one way or another (either being too verbose, or lecturing me about how what I'm asking for is evil and bad).
Can you give an example? It's a bit entertaining to read this given all the years long (and still ongoing) posturing about LLM sycophancy (not that all of these couldn't be true at the same time).
That seems about right. 4.8 is like in between 4.6 and 5 in terms of capability and language and they are all pretty close honestly. I just default to using Opus 5 for a coding agent that I don't interact with and I like driving with Opus 4.6 or Fable. Fable thinks too much though. Fable is like that engineer on your team that will over-engineer the shit out of something if you let them. Fable was like, "Here are 34 yaks, which shall we shave first <hands rubbing together>" and I was like can't we just... write the script first and then decide of any of these poor yaks need shaving?
The other thing I forgot to mention about Opus 5 is that at least out of the gate, it seemed very intent on spinning up agents and obliterating my token budget. It was noticeably more token hungry than 4.8. It would make sense for them to intend this behavior.
Also on a dark reader? It says at the bottom, but gets dimmed out pretty seriously with dark reading. It's a 7 day average business spend, relative to June 1st (2025 presumably) indexed at 100.
Might as well not be - I routinely get rate limited in a single review session.
I've honestly stopped using CC and moved to codex. Sol has it's warts but I've never once hit limits on a 100€ and I get a similar level of performance for what I'm doing.
I wouldn't mind bumping Fable to 200$ plan if it was actually better but between the insane caps, reverting to opus/sonnet randomly and having similar perf as OAI - I'm done with it.
Next step is to put 100$ into open router and try some western hosted open models with Pi when OAI starts pulling up prices.
I did not know that - but what happens then I get denied and I plead with the model that my usecase doesn't violate its TOS ? Edit prompt till I get it to pass ?
Reverting is annoying but being flagged for security questions while I'm doing code review is insane.
Yes, you get to re-try it until it might pass. Each attempt uses tokens at Fable level though. I once burned through a 4 hour window ($20 plan) in 4 attempts to re-frame the request.
You want to disable "Switch models when a message is flagged"[1]
Yeah for me Fable works great on the $100 plan. The problem with Fable is you will hit the weekly limit. The $200 plan does not give you 4x or even 2x the weekly limit. For me i only get about 1.3x weekly quota, so I don’t think it’s worth it.
I still havent used it much, and they crippled Opus for what feels like no reason. Some suspect that Opus knowing theres a better model degrades it some. Shame.
I'm on the $200 plan (work pays) and I also have the $20 OpenAI plan (I pay) and keep a balance on OpenRouter.
There is nothing as good as Fable, not even close.
I recently had it run a 18 hour autonomous rebuild of a project (moving from Spark to Pandas for performance/data size trade off issues).
It orchestrated Opus sub-agents flawlessly for 18 hours. It even did a great job of managing the number of agents to keep them within the 5 hour budgets (I think I had to restart it twice).
After 18 hours I ran a /simplify, /code-review, /simplify cycle which went for another 6 hours.
2 billion tokens (mix of Opus and Fable), 24 hours of continuous coding and a bug free outcome. It would have cost $2000 at API prices and worth every cent.
Fable's ability to keep other models on track while working on these long horizon goals is so much better than anything else.
Far from neutered, I've never had a cyber refusal, and Fable's English is actually readable (unlike Opus 5).
As an aside: while I hate reading Opus 5 English it still is a noticeably better model than Sol in my experience.
But I could handle losing Opus5 is I got Sol instead. But there is nothing even close to Fable.
I like this story. How did you verify the output? How big is the codebase? Why it took 18 hours? Could you implement it with a small local agent and breaking the task down yourself in two days (i know it sounds like a loaded question, it is not).
I think the rewrites are the main story for llms in code (hot take). Writing greenfield code at the seams also something which might work well.
I think you need to spend more time with Sol. If you think there is nothing even close to as good as Fable - my guess is you haven’t spent as much time getting as familiar with working with those models as you have with Claude’s.
Codex is more token efficient and tends to get better results than Fable with less need for extreme token burning shenanigans like 18 hours of subagents.
I spend 10+ hours a day in both agents, typically side by side. I often have them do direct “bakeoffs” from identical prompts in separate work trees. Most of the time, Sol’s work is better than Fable’s. Not always. It’s situational. But it’s certainly not the case that Fable is in a league of its own or anything.
the sentiment is negative and justified.
anthropic nanny states what you can do.
in your instance, anthropic may decide, arbitrarily, to stop 'autonomous rebuilds / refactors and ports' because they could pose some alignment/rights/etc risk to whatever slop their philosophers dream up while they're out eating $200 avocado toasts. then you can't do the thing anymore.
fable is good, absolutely. agree it roasts Sol which is, comparatively, a little receipt-hunting jack**
but now imagine being an enterprise, and having another organization not only taking your workflows and baking it into your models, but then deciding they can arbitrarily cut you off.
when you can instead own your data, use an agnostic provider, and get better results (through model combinations), it will take 1-2 quarters to figure it out.
the main reason anthropic is killing it is because they really do understand the enterprise development experience and lifecycle and have built products and have a sales-team that can deliver.
business-model and vibes-wise they have lost all goodwill in the past 6 months, and that momentum will be quite hard to regain.
Mythos? That's only for super special corporations, you need not apply. Fable? Can't even look at it wrong without running into cybersecurity lockouts. And even if you get past this, it's limited to <50% usage. It took me a month to complete a Fable review on my project.
So stingy. Even with OpenAI's recent usage troubles, they're still so much better than Anthropic it's not even funny. Re-ran the code review with Sol as a benchmark and it turns out Sol's performance is within 70%-90% of Fable's. Anthropic's still got the best model, but what does it matter if I can barely use it?
50% usage plus there seems to be a pretty big metering multiplier still. Do a relatively in-depth review of 3k LoC with Fable xhigh and poof, there goes 5% of the weekly Fable allowance. If I use their first party code-review skill that spawns a bunch of subagents—well just forget about it.
Non-LLM user here. Why? Apart from the ecological issues, I'm very uncomfortable giving my organisation's crown jewels to {random_internet__corp}. Look at the lengths they go to for training data - 10M for Spirit's call logs? Destroying millions of obscure books to scan them? They make meth-heads look scrupulous.
Your prompts, especially if they contain your entire codebase, are _way_ more data-rich than an airline phone call or 1920's novel. So they _are_ going to train on them, no matter how many checkboxes you tick to stop them. Which is both a commercial risk and a massive new attack surface.
It's not just software developers; if law firms aren't controlling any public LLM prompts _very_ carefully, they can expect some meaty client confidentiality suits. In fact any organisation in a competitive environment should be worried: Acme Bolts: "Write me a presentation for Zoom Construction". Beta Bolts: "Is Acme Bolts pitching to Zoom Construction?"
Maybe this is part of why OpenAI and Anthropic are finding demand softer than they would like. And why open-weight models that organisations can run on exclusive hardware are thriving.
For many coders including myself, LLM based coding agents work well enough to be useful, and in some cases worth paying for.
What I don't see is vast areas of industry finding $10s to $100s of billions of value in LLMs. There's no lint or compiler that can check for correctly constructed contracts. So LLMs, which should be useful to law firms, incur a lot more manual checking of their work than coding agents.
Less formal document production in other industries is likely to have less structure. That might not matter in some settings but I'm having trouble thinking of an example off the top of my head.
While there's no linter for writing contracts, my experience (as a commercial lawyer) is that frontier LLMs are far better and error checking and far quicker at writing than the average senior lawyer. The main thing holding back further deployment (in my jurisdiction) are concerns around data residency, privilege and how fundamentally it will break an industry that is so heavily reliant on time based billing.
Using them for coding makes it easy to self check its work (assuming those pieces of work are "verifiable").
As proofreading will still need to happen, what do you think the appetite for lawyers is to do this kind of work? Do you think it will drive fees down significantly? Empower younger lawyers at firms who probably are the ones doing this checking for the partners? (Or will that just create a further divide).
I'm genuinely asking as I am not in law but all my family is and it's nice to see someone here that's thought about the impacts in that space.
I compared insurance quotes last week. Needed cover for 2 brands, 1 company. Opus jumps up and down saying both brands need listing on the policy schedule. Human broker said not.
I told opus and it's the usual "thanks you're right" bollocks because it bothered to read in more detail and found that all business activities are covered.
I have a family member that is an attorney in housing law. She claimed that LLMs are not particularly useful for her work. If she asked a simple question like, "Find all the <insert specific housing laws> for all 50 states," then she still has to go and check every single one of the laws. Since the legislature is modified so often, she cannot look at, say, Maryland's law and know if it the LLM output was the 1990, 2014, 2018, or 2026 version of the codified law. In order to fact check the law, she has to look it up, and by that point in time, she has the answer she did the work of the LLM.
This is my experience if you need the data to be accurate often it isn't and you only know that because you had to check because it was important. Things not important you never check.
But translations don't require anything close to SOTA level models. Translations will be high volume, low margin transactions. That will not save Anthropic.
International trade is a giant sector already, and potentially much bigger than that, now that LLM assisted translation makes it easy to offer your products and services in any language, or make complicated and sensitive deals without a common language.
Hackers always rage and down vote every time I mention this, because they are unable to see beyond their small world. Why didn't they learn that their part of the internet is 0,000000001% of what the world uses the internet for today. It's going to be the same with LLMs. Programming and hacker stuff is going to be 0,0000000000000000000000000000001% of what the world uses AI for. But translation is going to be in the top 5 of use cases.
Product marketing and contracts will be a 'common use case' we think of, but it will be 0.000001% of tokens consumed.
A developer using sub-agents will consume more tokens 1 Day than a marketing manager will consume 1 Month, easily.
Unless there is something inherently automated about the nature of the AI, it will be a tiny % use case.
Even a lawyer, using AI daily for contracts - that will be relatively light use. They'll make more use doing legal research etc.
Developers and Automation are the 'primary' uses cases for AI, and in the future, we'll start to see AI integrated into Apps - that will be 85% of tokens consumed.
Yes - once translation becomes realtime, and we have our Star Trek Universal Translators, then translation will become more visible, but even by then, a relatively small part of overall consumption, even if it's more highly visible.
I think quite cheap models will probably work with translation (I think for many use cases Google's now fairly old tech works well enough and is free?).
The main problem is I think you're assuming because the current translation market is large (I'm just going to assume it's ~100B in size just from a cursory search), then it will remain large with LLMs. If LLMs are much cheaper than humans, even with a lot of growth in translation volume the total spend may not compensate for it (again most of the volume will probably be using almost free models?). Another is assuming that because something is valuable you can charge a lot for it. Like, oxygen from air is extremely valuable to us. If oxygen somehow depleted we would die almost instantly. It does not mean everyone goes around purchasing oxygen or even less that you can charge absurd amounts for it.
Transformer models are quickly becoming a commodity, and I suspect in time we'll all be running them locally. Even now, you can run something pretty useful on a 16 GB graphics card, and I suspect a decade into the future, entry level hardware will be running better models than high-end graphics cards can run now, as entry-level hardware gets better and models get more efficient.
It doesn't mean hosted frontier models wont exist, they'll just be rare. It's no different than any other commodity market, for example most cars are cheap commodity models, with rare individuals buying expensive luxury cars and businesses buying expensive trucks and specialized equipment.
>and I suspect a decade into the future, entry level hardware will be running better models than high-end graphics cards can run now, as entry-level hardware gets better and models get more efficient.
I doubt it. The play seems to be: lock what was once commodity compute up into datacenters depriving us regular folk of it, then sell it back to us on subscription. Even if my #NeverSubscribe movement succeeds, all that misdirected hardware [into datacenters] won't likely be practical for home use.
> The play seems to be: lock what was once commodity compute up into datacenters depriving us regular folk of it, then sell it back to us on subscription.
I had not considered that as a possibility. It is somewhat dark and unlikely, imho, but a real possibility.
I am more inclined to believe that the fab capacity will grow over time and "commodity" compute will be available to us all again.
However, I can also imagine going back to the 60s era "hyper verticalized" mainframes, in which case, the frontier labs might not just be producing models, but chips and an ecosystem around themselves and their suppliers/customers.
Well nobody thought you could write software that helps tightly coordinate processes happening simultaneously from millions upon millions of nodes on nearly every corner of the world, but here we are. When industry expands further off earth, we will need more complex and intricate software to coordinate its movements, why wouldn’t our systems become more powerful. If you are hopeful for humanity than you must expect the scale of industrial necessity to only ever increase alongside the imagination and capacity of its people’s.
My company still hasn’t been able to deploy wide access to Fable because it’s not available on a ZDR basis. This wasn’t mentioned in the article but I imagine this factor is not irrelevant.
It’s far from clear to me that this is directly connected to the no-ZDR requirement. I’m heavily involved in this stuff with my company and I’ve never heard that the lack of ZDR fable is a Trump admin thing.
Of course I echo the universal: Opus 5 sucks. But what also sucks is still using 4.8. Are you all seeing this? It's like the older models got dumber just before Fable and Opus 5 were coming out? I've heard the theory that it's because 4.8 is getting put on older hardware? And I imagine that just ratchets down reasoning time then possibly?
Sometimes I'm just trying to sus out if I'm truly seeing things these days or going a little nuts :)
I have noticed very serious degradation in performance from Anthropic. I've switched away from Opus 5, but 4.8 is still much worse than how it was before Fable came out. It's constantly making what seems like obvious mistakes... I point them out, it's constantly apologizing.
I don't know why...
Is it A/B testing?
Is it load shedding?
Is it because I'm in Canada?
Is it because I'm not on the Claude Max plan?
Is it because I'm not paying via API?
Is it because I'm not paying via Bedrock?
Is it because the U.S. is worried people are distilling?
Is it because the U.S. wants to keep the top capability to themselves?
I think open models are the future. Anthropic is killing their reputation so fast. If they don't come clean I think they're cooked.
This is yet another reason why I think local models will win in the future. They're almost certainly A/B testing all sorts of opaque stuff that people have no clue about, hence the various 'How's Claude doing this session?' popups.
So what you are paying for may vary on a day by day basis, which is quite undesirable, even if their main goal is simply to make a better model. When it comes to a tool, I'd rather have consistent mediocrity than instability.
I definitely agree. I honestly cannot do tasks that require even minor complexity. Opus 5 keeps forgetting things in context as well and coding conventions. Really cannot build with CC without Fable.
I'm not an expert by any means whatsoever but deploying models is not a straightforward task. There's a lot of levers to pull and I bet when models get "downgraded" to older hardware they do so WITHOUT the same stringent quality control of the output as they do when they release it.
I don't think it's something deliberately malicious like planned obsolescence but it's more like startup culture of "just make it fit in this sprint".
For me, Opus 5 mostly sucked because of its incomprehensible writing style. Having the system prompt focus on writing in terms that are easier to understand helped a bit
For _most_ day to day knowledge work (writing, excel, filling forms, coding) current SOTA models are good enough. There are diminishing marginal returns from paying more in my opinion.
If you are disproving Jacobian Conjecture it makes sense to be on SOTA, but for writing Golang and Typescript, faster sol/fable/opus class models are imo more likely to get user interest than the latest frontier.
This is really it for me as well. At the heights of complexity AI can do magical things. But really a lot of the time I just want it to do mundane things right. And currently it just cannot. It writes garbage text, consistently ignores something you have told it, makes mistakes a human makes once but the AI remains uncorrectable.
Dijkstra in the Foolishness of Natural Language Programming
[...] the "naturalness" with which we use our native tongues boils down to the ease with which we can use them for making statements the nonsense of which is not obvious. It may be illuminating to try to imagine what would have happened if, right from the start our native tongue would have been the only vehicle for the input into and the output from our information processing equipment. My considered guess is that history would, in a sense, have repeated itself, and that computer science would consist mainly of the indeed black art how to bootstrap from there to a sufficiently well-defined formal system.
Imagine if only we had languages at our fingertips whose explicit purpose was to precisely and unambiguously tell a machine what to do!
I’m not sure if the underlying data is counting subscription use for Fable, which is where a lot of people are using it because token pricing is very expensive. I wouldn’t be surprised if this was counting enterprise token usage only. As rich as enterprise customers are, they’re not exactly willing to double the cost of SWE salaries on tokens.
Either way, a model used to solve the top 10% of problems that people use AI to solve for, being used 10% of the time … seems like it’s in a decent place.
I find Fable indispensable, and measurably better than alternative models, for complex feature development in an existing codebase. It’s the closest thing I’ve seen to nearly one-shotting features. Even still, I only use it for the hardest features and Opus 5 does a good enough job on the rest.
Very much agree on Fable. Over the past month or so it has shown to be the only Anthropic model that can understand a largish dbt model codebase. Opus 5 gets almost everything wrong. (High reasoning on both)
More importantly other way cheaper models can do what Opus 5 does. So you can pay for Claude to use Fable 5 exclusively for harder stuff and planning, then get the same value you'd otherwise get from switching back to Opus by using other cheap LLMs for day-to-day coding tasks.
Yeah as soon as CFOs realized AI was racing to become one of the most expensive line items along with salaries and AWS bills, they started cracking down on the most expensive ones.
The US government basically told them they can't sell "Fable" and so they aren't. That's probably 90% of the story.
Once the government stepped in, their 5th generation was effectively killed. They either had to lock it up (Mythos), neuter it (Fable), or leave people with the perception they were overpaying for a weaker model (Opus 5).
Hopefully Anthropic has learned to anticipate this risk and has a plan for rollout of their next model that plans for capricious ad-hoc regulation.
>They either had to lock it up (Mythos), neuter it (Fable), or leave people with the perception they were overpaying for a weaker model (Opus 5).
Option C: they could "get hacked and the models exfiltrated" plausibly. Then the horses will be out and there would be no more point for the USGov's barn door.
They tweaked Fable to appease the USG, but it clearly refuses/downshifts more requests than e.g. Opus 5. (I have run into Fable refusals even doing really normal stuff like fixing bugs reported by static security scanners like Brakeman. Not downshifts, outright refusals.)
So the net result is they put "Fable" back on the market, but it was perceived to be a worse overall experience, at much higher cost. (Remember all this happened before most users had really even had a chance to put Fable through its paces.)
That's where we are today: Mythos is locked, Fable is neutered. They probably can't fix this until they are ready to release something that's not called Fable, or that they can say has some key architectural differences to Fable. The sword of capricious regulation is always going to be hanging over Fable.
What did they think was going to happen when they were actively doomsday fearmongering around the release of their own model? I assume somehow they thought this would lead to a moat with them being left safely inside the castle, but this response, and Reagan's 9 Words, were always infinitely more likely. It was a demonstration of a child-like level understanding of how regulatory capture works.
The PR from the US frontier labs has been poor to terrible from the beginning.
Among other things, "we have stolen the collected works of your culture, now help us grow so that you can all lose your livelihoods and become our serfs" has to be one of the absolute worst marketing approaches in history.
The government seems like it's in a rough spot. If they let Mythos out, they seem worried people could use it to mass-hack the internet. China seems to not care so much about this and they're right behind. I don't really know what the answer is.
I wonder how much of it is fears over "mass-hacking the internet", and how much of it is fears over the model discovering various NSA/CIA/etc "tailored access operations", and other deliberate side-channels / vulnerabilities / etc?
(Or perhaps vulns that NSA/etc discovered and has been keeping it private; as they're known to do).
I feel there's a lot that's unaccounted for, and the whole "AWS team reports a 'jailbreak' that is just 'review this codebase'" story doesn't add up.
I wonder if there were some parallel construction going on, and if at the same time, the NSA started losing the exploits they had because it was getting patched.
I think the latter is likely their primary concern. Certain models making hackers more effective doesn’t change that black hat hacking will still be illegal and that’s always been the prime deterrent against capable hackers.
While there is some fallout likely with much more effective hacking being easily accessible, I’m sure govt analysts (unless they were let go) have their own prediction models telling them it’s inevitable that this technology eventually makes it to everyone they don’t want having it, what with China seemingly releasing every progress they make openly. Which makes me think that they’re preparing for that inevitability by hardening the govt systems currently in place and/or by burying the secrets they want to keep hidden deeper underground.
Given the history of the US and this particular administration, I feel burying things deeper is a greater priority.
China uses the strategy of letting dangerous technologies loose which causes disruption in the short term but makes people do the right thing like secure their software. The US by comparison gives me the impression of wanting to leave the internet vulnerable by not making Mythos public so that only the US government can use Mythos to gain access to whatever system they like, which is the same thing the pegasus software does
Whether or not the government chooses to regulate the space, capricious after-the-fact regulation is the worst of all possible worlds. The ~equivalent models from OpenAI did not get the same treatment (favoritism?), and it's not clear the government has produced even rough guidelines about how to be compliant going forward.
Model prep costs far too much money to operate under this kind of regulatory regime.
Fable is just way too expensive and limited compared to GPT 5.6 Sol, and the only task that requires that level of intelligence is frontier scientific research. I use GPT/Codex primarily for coding and usually keep Claude on Sonnet 5 most of the time as I use Claude primarily to debug/brainstorm/make frontends as a supplement to GPT.
I do not agree. Fable is the only model I can leave unattended and give me results for part of my work which is just devops related tasking.
I can hand hold opus but I would rather just ask fable to do it and give me the result that I review and works. Opus will waste tokens and still require me to help nudge it in the right directions.
I think the next gen models from china will put us in a spot that the cost can plummet and I won’t need the Sota from anthropic
I second the parent comment. 5.6 sol xhigh is not only better than fable I can also run it forever without worrying about limits. The frontend has gotten much better too with the plugins that come with codex.
Agreed. I have the highest individual plan for both. I run out of Fable credits midweek, while I usually have some credit with Chatgpt left despite it having to carry Fables load for the second half of the week.
Also, Sol doesn't refuse constantly and speaks like an engineer rather than a deranged academic.
Opus 5 is legitimately terrible and can't or won't follow instructions. It is of negative utility and does more harm than good to my codebase.
In a frontier scientific research environment, funding is often limited, so personal subscriptions are more common. Fable can hit a 5-hour usage limit on the Max subscription tier before it finishes a single complex math prompt. Most of the time you get one or two if you reduce context and question complexity. Most of my colleagues are cancelling their subscriptions because of that, and just using Sol instead, which is virtually unlimited on the Pro tier.
Hmm. Would you mind sharing an example of a complex math prompt that you would use? Because I found that even Sonnet can solve fairly complex math problems fairly easily if you give it the right tools, so I'd like to give it a shot myself if you'd like.
> the only task that requires that level of intelligence is frontier scientific research
Even fable is way below human level at many tasks. If fable were as good as you claim, all computer jobs other than "frontier scientific research" would have been replaced by now.
Most of the benchmarks have exceeded their usefulness. Opus 5 beats fable 5 on many of them. Anyone who has used both models will notice immediately that this doesn't translate to the real world. Opus 5 is nothing short of a regression from Opus 4.8. Fable is genuinely a great model so long as you don't trigger a guard rail and it downgrades.
Sol in my experience isn't significantly different than fable ignoring that Sol burns usage 10x faster but the end result is hard to differentiate.
GLM 5.3 is a hair behind these two.
An anecdote but not an original one from the people I talk to.
Sol routinely catches stuff in review that fable misses for me. It's impossible to compare them meaningfully because it's a complete dice roll - but in practice using both in my projects I can get work done with both, and Opus 5 is far more tedious.
But Fable security false positives and pricing just make it not worth compared to Sol IMO.
Yeah, it's pretty much apples to oranges, and I don't consider GPT and Claude to be interchangeable at all. From my anecdotal experience, GPTs generally codes more creatively and verbosely but Claudes tend to code more carefully and precisely, so the result is that GPTs generally finds more creative solutions to problems but also writes buggier code, which is why I converged on the setup of GPT/Codex for implementation and Claude for debugging, which feels more like a force multiplier than using each model individually.
There is a difference between needing frontier capability because one is solving a truly open ended problem, and needing a reliable workhorse model to do something well understood. Local models (like Qwen 3.8 27B) have gotten so good that they can do all routine tasks at a fraction of the cost of frontier models.
Anthropic’s issue is churn because of the peak verbosity vomit coming out of Opus 5/Mythos/Fable. What the hell did they train it on. The sane model is still Opus 4.6.
- Opus 4.6 was the last Opus generation that got a lot of use by Anthropic's own employees
- After that they primarily used Mythos internally
- 4.7, 4.8 and 5 were RLAIFd by Mythos "teachers"
- Hence why 4.6 is the last Opus gen who doesn't report back like a robot wanting to cover every potential hole another AI system would've spotted and criticized
- Hence why coding style in Opus 5 also gets criticized, not only behavior in CC
Commented elsewhere, I still use Opus 4.6 because it is the only model that feels decent to interact with. 4.8 is decent and some times smarter but you can see it trending towards Opus 5 levels of nonsense. I use Opus 5 when I don't need to interact. Fable or Opus 4.6 are the only Anthropic models I like interacting with ATM.
In my opinion, the big issue with Fable is that Claude Code cannot use it properly. I know, that sounds weird, but I've had Fable run down the wrong lane (and never stop) or give up and claim that something was impossible so many times (until I pointed at a GitHub repo that solves the "impossible" issue).
But a while ago, I had access to a Fable harness that just never gives up. And that verifies itself. It burned $100 in API tokens in 15 minutes ... but it succeeded for all the prompts where Fable + Claude had failed.
And I believe that's a real issue for Anthropic. Fable+Claude is not too expensive thanks to the subscription, but Claude severely nerfs Fable. To save money, I guess. Fable API + Custom Harness is a different class, it's so much better. But API tokens are so expensive, you're cheaper off hiring a freelancer.
It all started when OpenAI began focusing on Codex and coding while sunsetting Sora, which helped free up a lot of compute and resources. 5.6 was the final nail in the coffin.
Great find. Related: we are building OpenAmer, a fully open-source agent that controls the actual desktop (files, browser, terminal via CDP), has persistent memory and A2A multi-agent swarms. Apache 2.0, runs local on Windows: github.com/openamer/openamer
GLM series has made it very practical to self host. If the new update for Deepseek flash holds up, I think it would be silly for some companies to not self host.
I wish Fable were as good as you make it sound. A plan created by Fable is good, but in my case, it always contain issues caught only when it's reviewed again (whether by itself, Opus, Sol etc.). That's (almost) not different from plans created by Sol, GLM 5.3 etc. The one thing where it's genuinely better is the front-end, but then again it's far from perfect, it just needs less iterations.
>> but in my case, it always contain issues caught only when it's reviewed again
Yes but those issues will be much less severe with Fable-written plans than those written by lesser models. I know this because my workflows at both my regular job and my startup involve multi-step agent reviews via codified adversarial review skills. Fable as a reviewer will frequently find blocker-level issues with plans written by GPT 5.6 Sol, and sometimes with Opus 5. The opposite almost never happens. In fact I cannot remember the last time it happened.
I have a sneaking suspicion that someone at Google may be making the same bet, looking at the faster and faster Flash models which provide acceptable results to a lot of people (outside of coding).
For Anthropic. But not for the AI industry at large.
Cheaper, more powerful AI will continue to expand the bubble. Projects will get more ambitious. Everyone will build out their own custom little software. Code diversity expands and requires even more AI.
The bubble is _not_ on models becoming more intelligent and solving arc-agi-999.
They are already good enough at what they mechanically are.
You have to use the right harness, right verifiers (automatic where possible, human where not), etc much much more specific than a generic one like claude code or codex, and it will also be able to work within constraints and be the "proposer" of an imaginary optimisation problem and an excellent one at that. But you have to frame the task at hand in that manner or maybe even reorganise the task you do itself so it is more amenable to being framed that way. If you use it this way, it is _already_ massively economically useful. But it will take many years for it to actually be usable in that way, since you need DC capacity to come up first which is few years away and also well, massive organisations that have to integrate these will usually take many years to do so.
It is also useful albeit less so in cases like general SWE, where you still need a human in a loop for non-verifiable requirements, and also in other general usecases where information retrieval is too intractable and you need to carefully use LLMs as a component of the overall system.
I am not saying Fable or whatever the biggest models are are useless - they will certainly be useful for tasks at the frontier of the day - which is today complex exploits and open math problems, and well, tomorrow it could be something in biotech. But this is not what the entire bet is on at all - just automating day to day drudge at the tens of thousands of massive companies and governments we all know and love is more than enough. With the right training data (which _also_ is a bottleneck and takes time) you could even automate certain processes entirely. Sure, if we get a crazy medical innovation and end up saving trillions in healthcare great, but that's just a bonus.
None of this is to say that I think there is zero sketchy financial engineering going on
The bubble is based on the promise that these LLMs will cure cancer and find the solution to global warming. The pragmatic users of these tools (like you seem to be) are enjoying the subsidized use of the tools right now, but it's not a sustainable business model
The issue is that, if there is a “good enough” point approximately here, it is only a matter of time before models become small and efficient enough not to need all those data centers. Though, it should be good for companies that sell computers (like Apple) rather than putting a toll booth in front of a pile of numbers.
I think highly complex work and “professional” work are basically completely orthogonal. You can have highly complex work you do as an amateur, where AI can be very useful. For example working through a difficult mathematical problem, building or contributing to an operating systems, or researching a highly technical topic for a hobby project such as microscopy, chip design or lithography.
And you can have extremely simple work that you nevertheless have to do as a professional. For example, drafting a routine customer email, summarizing a meeting, formatting a report, filling in standard documentation, or making a trivial code change.
So I don't think the $15k workstation / professional camera analogy really holds. Those are specialized tools whose capabilities are mostly useful within a fairly narrow domain. A general-purpose AI can be useful across thousands of completely unrelated tasks, including one-off problems encountered by ordinary people.
As a small background, I have a local server and I've been trying out different models with different inference engines, quants, configurations etc... I'm also using Opus and Sol at work consistently. I've used AI since the first wave, first as a toy, then as a highly specific tool, last 6+ months as the primary LoC generator.
This is the first time I've felt, and I use the word *felt* since I don't have a suite of benchmarks or any sort of material approach towards comparing models, that Opus has declined in quality compared to before. Primarily I think its powers of deduction and understanding, even on xhigh, have become much worse. Before, being vague and providing a simple prompt would be enough, it could deduce and expand the details it needed, plus ask you clarifying questions, now this is no longer the case. A concrete, personal example, for a personal project, I've asked it to setup ssl over local IP. I didn't go into too much detail in the prompt as there are many approaches it could take and I didn't care too much to choose. It did horrible. The first thing it did was say the best lightweight approach is to add a reverse proxy. I'm like ok, makes sense. Then after asking it to proceed, it went and added a bunch of config to my golang service and didn't even setup a reverse proxy even when it said that is the way to go. It even said it didn't set it up lol. Then after I told it to do so it failed building the config in a way it was asked of it (support LAN IP and tailscale IP). Etc etc...
When Fable came out it was huge, the benchmarks told the story, and the story mostly matched the experience. It felt, again, intentionally saying felt, like it was miles ahead. Now benchmarks say that there are many models that are close, but in actual use Fable still *feels* much better. I think benchmaxxing the new open weights models is ruining the value of benchmarks, if they ever had any. When you actually put them to the test you see 500k tokens of reasoning with "Actually..." and "Wait..." in every third paragraph of their reasoning trace.
The price for Fable is definitely too much for any personal use now that it's no longer included in the subscription, and GLM 5.2, Deepseek Flash and Qwen 3.8 served locally or via cloud provide a lot, requiring a bit more babysitting though. Considering the price of Fable, my 5k USD Epyc server would pay itself off in less than a year if I used Fable or Opus in the same manner so at least for me the decision seems easy. And considering the point I'm poorly trying to make, that Opus doesn't feel like frontier anymore, this is probably the last month of my Claude subscription.
> When you actually put them to the test you see 500k tokens of reasoning with "Actually..." and "Wait..." in every third paragraph of their reasoning trace.
I've wondered if this is part of why we don't see the reasoning traces for Anthropic's models before -- Open models might just be accurately surfacing how the sausage is made.
Hasn't it always been the premise that intelligence would get cheaper? To me, on the enterprise side, it seems like firms are finally getting the memo that, whether you are locked into the Ant/OAI ecosystem or not, you don't need the smartest, most expensive model to do every single task. This is a good thing for overall adoption. Whether that trickles down into regular user behavior, especially with subscription pricing, remains to be seen; even though I intellectually know I don't need Sol for a simple refactor, I am sometimes hesitant to choose Luna/Terra, as it's hard to accept using something positioned, even implicitly, as 'worse'. Remembering that the smaller models tend to be faster is what usually pushes me over the edge.
Anthropic in particular is much more compute-constrained than OpenAI and SpaceXAI and has relied on partnerships to provide inference. This reality factors into their pricing and usage limits (they started 'adjusting' the 5-hour limits during peak hours, and it certainly wasn't an upward adjustment). Accordingly, this is presumably what Anthropic wants, given they develop and release the lower-end models, suggest users use them in various nudges within their product, position the bigger/more expensive models as "For the most complex tasks" in their UIs, and so on.
> whether you are locked into the Ant/OAI ecosystem or not
I think the problem (for Ant/OAI) is that there is no sensible lockin or moat. LLMs are essentially interchangeable and stuff like a harness doesn't offer enough value on its own for someone to be locked into using one of them.
Now with the onslaught of the Chinese models that offer almost the same quality for much less money they have a very serious problem on how to proceed. Investors now might be looking through rose tinted glasses but their patience has its limits.
Agreed 100% for the consumer case: an empty chatbox is just about the least sticky surface I could ever imagine. I saw a mobile interstitial ad for Kimi recently whose hook was basically "Tired of paying for expensive ChatGPT? Download the Kimi app, it's the same thing but cheaper". I myself bounce between token subscriptions like no one's business and use Pi/OMP for maximum model flexibility when coding (and it's a few env variables or lines of (TO|YA)ML|JSON to switch providers in Codex, Grok Build, CC). I even self-host and try to use OpenWebUI + CLIProxyAPI when I can for all my chats.
Enterprise is a whole different ballgame IMO with countless technical, compliance, and employee adoption considerations that add friction to switching. It's also where both Anthropic and (as of last week) OpenAI get the bulk of their revenue, and, incidentally, the venue where US Government regulations on Chinese models would have the most impact.
Fable is not used because it has extra retention requirement that corps can't sign off so it stays disabled for everybody in many cases.
It's also not winning on day to day work against Opus 5, which is simply available as there are no extra retention requirements and no extra paper work to do with legal.
Devs also don't like that Fable refuses to work on half of their prompts.
Then comes better pricing and on-par capabilities from competition.
Like 90% of consumer AI use is stuff like "please write a summary of this PDF for me", or "Find a cheaper version of this product online", eventually the returns on model intelligence taper off for these kinds of tasks. However for the kind of frontier tasks they are testing the models on now, like math, science, etc. the marginal returns on increased intelligence are huge.
1. They acted as if they were so far ahead capability wise that they could stop listening to their users.
That is basically it. It is an extremely common belief that Opus 5 acts like a condescending wannabe-thought-leader, yet Anthropic's response to this is largely been "You're using it wrong. Try deleting all your config files".
They released a "concise" output format, but pretty much swept it under the rug, despite it being one of the loudest complaints about their models. It also, generally, does not work as described.
Not only did Opus 5 get much more difficult to work with, but it got, from a customer view, significantly slower. The token rate might be the same, but if every interaction takes 50% more tokens, it's 50% slower.
All this time OpenAI has released a slew of new models, lowered the price on them (which, bluntly, 80% of user count don't actually care about since they're on subscriptions), increased subscription capacity, and increased response speed.
It's not about cost. It's about Anthropic being the frontier-lab version of the marathon runner who decides to celebrate to early, and then loses the race.
Opus 5 is by far the best model for everything that matters to me. But it’s just too expensive. Terra 5.6 is bearable for everyday tasks, so now it’s my default.
I have yet to even try Fable or Opus 5. Just looking at the hype they put out prior to the release, then the whole way of releasing these models as well as the pricing just puts me off to get used to it and then needing it. And this while so many good models came out without any hype, botched releases and significantly cheaper. Anthropic really shot themselves in the foot. I went from using sonnet and opus models for 50-60% of my daily token usage to 10-20%
The rate of model improvement has slowed, and may not recover.
It's unclear if Mythos2 or 3 or whatever they're calling their next model will be an improvement for most common enterprise use cases.
LLMs can't solve basic things (writing non-slop documents, understanding context without massive handholding) and for coding other models are quickly becoming 'good enough' without the same cost and nannying.
That's why Anthropic is 'stealing' workflows.
But it turns out it's much harder to push adoption when your users don't really want to use your product.
Code was a unique use case where the code luddites were loud but a minority - most people don't want to update 300 cases of variables across their code base for a name change. Most don't want to write unit tests.
There are a few use cases where that will happen (law is next, maybe quant finance) - but otherwise most companies are throwing money into a pit and getting 0 return.
It's a very interesting race and state of affairs, but Kimi K3 and likely the next DeepSeek models will put the high price token affair to rest.
Unless of course, mythos / next model really does solve some universally applicable problem that people want it it to do.
Perhaps nerfing the cap out of your best model for press attention and hosting valuable features like thought traces isn’t such a great business model?
Are we at the stage yet where a super-strong model that needs oodles of safety protections to stop it outright hacking you is strictly worse for day-to-day tasks than a much cheaper model that's simply not competent enough to be dangerous?
If you are still investing in these companies or plan to in the IPO, the financial ruin you experience is your own doing, you are ignoring every sign that this isn’t going to work out. None of these numbers make sense and point to AI being a commodity with razor thin margins and a race to the bottom. Would you invest heavily in a toilet paper company that took massive amounts of power to produce each version that is 1% softer or stronger every six months?
A model is just another thing to plugin to a harness. I don't give it much more thought than that. If developers are still caught up on Claude Code, or Codex that's just not a long term thing. It's best to develop workflows locally and in the cloud with open harnesses. I know this will be the future because that's how it worked on every other system that developers use.
Sure there are Microslop and Oracle db users but most of the world we live in is Postgres and Linux. That's why I think most companies will run llm's like that.
That’s not true. Before AI, I have been using Jetbrains IDEs as far as I can remember. Also have been using MacBooks for work since my first job. You don’t have to generalise everything. If a particular specialised tool is good at its job just use it instead of re-inventing the wheel
> good at its job just use it instead of re-inventing the wheel
Exactly why is everyone reinventing a harness every month. There will be Microslop / Oracle harnesses and there will be 1 or 2 open source ones that win.
I have a $20/mo subscription and have found myself getting limited every few minutes of late. I'm just building a simple website, but after about 15-20 mins of back-and-forth, it tells me I'm tapped out for another 5 hours.
This wouldn't be so bad if it didn't make mistakes periodically, especially when it's about to tap out. I get the sense that if I upgraded to the $200 subscription it would get me a lot more usage, but it would still run into these issues anytime I sat down to work for a few hours.
I'm just using medium effort, so it's not like I'm on high all the time.
I have 200MAX subscription for Anthropic and a RTX 5070TI 16gb GPU. I only use Claude.AI web chat and I only use Opus4.7. I build my prompts on prem with Qwen3.8:latest and copy paste them in claude.ai. I never type anything into claude.ai I am merely a copy pasting monkey. If I need to adjust something I do it from the the on prem prompt. I download code bundles from Claude.AI and push them to my git repo I host. I instruct claude to not do any testing and only let me do testing on my machine so i am not wasting claude resources. The point is to minimize the claude agents from doing anything but specifically doing coding. I can do this for 12 hours a day and reach about 80% of my weekly usage. Once I get to 3 hours left on my weekly I start a fresh chat session with Fable5 and have it do blind code reviews on everything. I only do this because I want to max out my weekly allotment and fable will get me to 20% in 3 hours on a 200MAX sub. Sometimes Fable5 does something interesting but for the most part my on prem static and dynamic code analysis have kept everything buttoned up.
With open models delivering close to frontier level capability, I really don't see how these labs can only focus on having best model to sustain their business. Although these model's usage will grow, I think and hope that the serving will be distributed among many players.
im using fable almost exclusively, i just buy a new 200$ license if i run out of capacity. its so much better, easily worth it given how much work I get done. I built a code, deploy, e2e test loop with my full aws infra, fable just implements linear tasks constantly, 5 at a time, tests the whole thing end to end
if people dont see why they need a model this smart, they probably arent using ai enough
People have realized you don't always need a Fable level model. Majority of my work is sufficient with ChatGPT Luna which has effectively unlimited usage on the $100 ChatGPT plan.
I am not doing awfully complex tasks though. I imagine a lot of other people are in a similar boat, either switching from Claude to ChatGPT or even just min-maxing DeepSeek V4 Flash 0731 or similar.
Claude will refuse to visit sites with robots.txt, make a graphic spoof of an iOS game, or even fill out an employee survey on my behalf. ChatGPT is always happy to oblige, no questions asked.
I can respect the guardrails - I also can see why OpenAI may not have much control over their models - but I need an AI who will do whatever I ask and not play judge and jury.
so whats the best and most efficient coding harness and against which model? what are folks doing to keep costs low? I spent $1000 just this weekend on my personal projects for sota Claude but I feel like I can probably get much more juice if I start looking elsewhere.
Yeah I mean if I run out of tokens every couple of hours and have to pause my work or shell out more money I’ll switch to other tools that don’t have this problem. Though they turned this down a bit it seems, I can work with Fable reasonably now and I enjoy it actually. I think they were just testing out how much they can raise the cost without users leaving when having the best model. I guess not much after all!
I see variations of this post all over X. Fable doesn't get much usage, since Opus 5 is like 99% as good at a lower price point. And perhaps more importantly also faster.
I am getting close to dumping Anthropic. I like their models in general, but boy I have hit their "F you, I ain't gonna help" too many times now on innocent things. Ain't nobody wanna deal with that. I have never gotten that from Antigravity and if I did I woupd tey Codex then go to Openrouter and leave US models behind.
I mean, who says "screw you" to requests to get 35+ year old vintage computers working? Claude, that is who. Its guard rails are so stupid. I hear people trying to do simple mailing list management hit it too.
Give me a break, Anthropic is expensive and their CEO is not the nicest guy around, and the games they play with other people's money/their API is not fair. I stopped using Anthropic and OpenAI once they started calling for regulation of open models. I have survived locally since LLama3-70 days and have been surviving fine. If I was to pay for cloud models, it definitely will not be Anthropic, Fable or not. From what I have read, the best AI model will not even comply with requests most of the time because it or/and Anthropic supposedly knows what's better and safe for you.
Honestly, this makes a lot of sense to me. Devs that are good at their job don't need the absolute best model for every task, and if a cheaper model gets the job done 95% as well, it's a pretty easy choice
People here won't believe this, but Anthropic will begin to decline after their IPO when everyone runs to good enough cheaper models to save on token spend.
I've always wondered why everyone flocks to SV's latest darling company. Have we not learned from our history of glorifying these SV darlings that turn hostile?
Most of the people pushing this are just hoping that they can cash out before the hype pops and financial gravity crashes the party. Sam Altman recently claiming that the singularity is here is so stupid on its face he should just be treated as what he is, a huckster.
None of this stuff ever made any sense on what it was being sold initially. It was always insulting that the media and business leaders tried to argue that the tech could replace entire call centers or vast swaths of entire industries.
People keep arguing, but it will or it has based on extrapolating certain, reasonable use cases. Klarna has shut up about replacing call centers with bots because Markov chains with memory only can do so much.
Some of us want to get work done and don’t feel the need to either glorify or hypothesize about what might happen.
I’m happy with Claude. If they become (bigger) jerks, I’ll switch to something else. I don’t ha e the energy to praise Anthropic today and I won’t have the energy to demonize them tomorrow. The emotional investment people have for/against these companies feels like celebrating or being offended by the weather.
A mixture of opportunistic edge-seeking, FUD, FOMO, novelty-seeking, the need to impress shareholders, the tendency of salespeople to believe other salespeople are telling the truth, the ever-present need to stay in front of relentless commodification, and pragmatic curiosity.
TBH I don't think any of that is unique to the IT industry's relationship with the Valley. Other technology-driven industries have a similar worship-ish relationship with a few rarified businesses. But the culture of the IPO exit accelerates all the most short-term motivations to do anything.
Reads like: brilliant Carnegie Mellon University computer science grad struggles to find job where he is not replaced by cheap, inferior Indian labor that still gets the job done, even if it takes marginally longer.
No shit we're all paying 清冲 Flash to do the grunt work. Turns out though, paying 清冲 Flash a few more cents does exactly what Ivy Wasp Pro Mythical does. Crazy how that works.
I am because Claude has figured out I am a biochemist and therefore even asking Fable what the weather is gets me bumped back to Opus, sometimes even Opus 4.8 instead of 5. Kimi? GLM 5.3? DeepSeek? No such problem.
I can literally open a new chat with just "Hello" and it gets bumped.
Memories break Fable 5 for me as well in chatbot. I ask Opus a lot of sec related stuff and now if I even type “hello” in chat it gets insta-downgraded to Opus 5.
Just cancelled my Claude subscription and used Ox Alpha Free through OpenCode (so that's one).
In my view, from the testing I did since Thursday, it's better than Fable. I had just finished a rather large task that Fable completed, including a /review and an /ultrareview.
Ox Alpha found bugs that Fable and Opus missed, and it continued to build things like a pro.
It does have issues with availability - but it's on a free promo right now. That also means that I don't know how much it would have cost if I had to pay API prices for it, which might not be cheaper than the subsidized small company / consumer usage, but for large companies paying API prices for either offering, the difference will be considerable.
I almost feel like there's some sort of paid campaign going on in Hacker News promoting the Chinese/open models. I feel like every day I'm hearing about how the frontier labs are dead but your experience is the same as mine. My company pays for Claude AND Codex but we never really use the open models for anything critical.
Doesn’t matter if your company pays for Claude, Anthropic and OpenAI valuations and expenditures commitment requires them to win the vast majority of the market to make economical sense. And the Chinese competition makes that very unlikely, to say the least. They likely won’t disappear fully but their “free” lunch as the AI darlings is done, on paper. Will be interesting to see how they adapt
Sure, I'm being "paid" by the fact that it costs pennies a day to use Deepseek and if I could afford the hardware then I could run it locally. Meanwhile, Anthropic is obviously a threat to open weight models and actively lobbies the US Government to have them banned or controlled.
So yeah, I and I guess others, are quite active in whatever little way we have available, to up vote new models and share stories.
It's absolutely a campaign. Every time anyone says anything about Claude, immediately and inorganically there's a bunch of people claiming to be biochemists who are constantly shut down by their work, people saying that they run out of tokens instantly even on Premium, people saying they get even better results on their GTX 4050, and people saying that Grok is better, or OpenAI is better. It's probably several different campaigns each run by different tranches of the competition. Reminds me of the good old days in which every criticism of bitcoin was immediately jumped on by nine or ten pretend Venezuelans who asserted that it was the only thing permitting their family to evade government currency controls.
I'm not a biochemist and I have been blocked by fable and opus for "cyber" just for doing things like asking it to ssh into one of my servers, look at a CVE, or do work in assembly. Been rejected to their cyber verification program 3 times already.
http://archive.today/ZLojz
Where Anthropic f'ed up was treating their monetization the way they treat model training. Turns out that success in experimentation is not transferrable.
They have tried to find the highest that the market pays for sota models; however, on the consumer side, this is just too confusing and unsettling:
"You can only use Fable for a week as a part of your plan" "Be ready! You have to start paying per token!" "Nevermind! we extended it for a couple more weeks" "Wait, now it's up to half your usage" "Ok, now its..."
Most people want to not care. We want our AI like electricity -- Kind of just there no matter how easy/hard is for the supply. You don't want your electricity company to be on the brink of cutting you off any second.
That's Anthropic. You don't feel they want to give you a dependable service for an, albeit premium, price. It's a constant bargaining game. That forces people to look beyond the walled garden. There, they find models that are fine... and without the shenanigans.
Yeah, I agree with this. The constant state of "...will the rug be pulled?!?" does discourage relying on it as a model and building a workflow on it. Anthropic used to just be a reliable thing you could play with. Now it's this constant source of anxiety.
It also didn't help that the government yanked it which adds another source of anxiety since OpenAI is on much better terms with the administration and the administration seems corrupt enough that they would mess with Anthropic if they got a big enough donation from OpenAI.
But anyway after Sol entered the picture, I don't think Anthropic can get away with this as much and I also think they're going to face a massive backlash from Max subscribers if they do end up ending the +50% promotion at the end of the month because Sol is a Fable peer and priced very competitively.
My wife's startup made the mistake of building her internal operations around Claude Team.
Then she hired a VA in the Philippines. Anthropic promptly banned her account without warning once the VA connected to the account. It took her weeks to get her account reinstated, at which point she had already moved on to OpenAI.
How many big tech companies let you talk to a human to get support. Automation is wonderful to cut cost for them but for the users being unable to get support is a horrible experience. But you cannot go elsewhere because they are the only player in town.
How can small companies with 1000x less money able to provide live support, but if you pay 20, 100, 200 dollars for a subscription you dont have a phone number to call ?
But Philippines is on Anthropics list of allowed countries?
It was flagged for account-sharing, or detecting a compromised account
How do you know there wasn't account sharing though?
I've seen some amazingly dodgy stuff when hiring people from south east Asia, sharing a paid account with friends worth a months rent there seems milquetoast in comparison.
There clearly was account sharing, someone logged into the wife's account from the Phillipines.
As soon as I started using Fable I was like, okay, this is probably as good a model as I will need for software engineering going forward. I still feel that way. I don’t need a better model, I need a faster Fable.
The thing I miss most about programming is flow, and the constant bouncing between terminal tabs sucks. I’d love to do one thing at a time, with Fable, quickly.
remember 4 year ago we use to : have stack overflow open, documentation, obscure forums plus other tabs.
An ide open with 20 tabs open each file a component, a class or an interface We also use to hold entire codebases in our brain.
You may or may not like agents mode. I also hate flipping tabs, but I enjoy using agent mode with well named sessions. I still stick with a single session until I must move to another, then I leave them around for a few days until I’m sure I won’t need to pick up where I left off again.
Command: claude agents
I don’t think they were complaining about literally flipping tabs but rather just needing to context switch so often. This doesn’t sound like it helps with that.
There is GPT 5.6 Sol on Cerebras if you want to try that experience for an ungodly sum of money (not getting into GPT 5.6 Sol vs Fable, but only one is available on Cerebras) for an 11x speedup.
> They have tried to find the highest that the market pays for sota models; however, on the consumer side, this is just too confusing and unsettling:
The consumer side cheap monthly plans exist for the same reason companies like Cloudflare and Vercel have a free tier: When it’s cheap and easy to get developers familiar with the tools, they will push their companies to pay the real money for those tools.
It’s a hard balance with LLM serving because you can’t really make it free. $20/month is close to free, but the $200/month plans are in a difficult place where they’re big enough that many small companies pay for $200/month plans for their employees and ignore the enterprise features you get with the full expensive arrangements. So the companies are continually adjusting the $20-$200 plans to keep them from being reliable options for businesses, which is where the real money is.
There’s a short sighted cheering on of the 3rd tier and lower companies offering lower rates, but we’re already seeing them ratchet up the pricing and keep larger models closed after they get market attention.
> That forces people to look beyond the walled garden.
Every time I get "you used your quota, come back in 3 hours, or 2 days" -> that is experimentation time with their competition, leading to changed service plans. When they said "claude -p" will be billed at API pricing even for plan users I moved my harness off claude. After I integrated codex, then it was never going to be a full claude project again.
What business encourages users to try their competition and adapt their usage to the competing products?
I agree they’ve done a bit too many pricing / usage promotions and A/B tests.
The period was also marked with many billing bugs, like spending people’s usage credits for included Fable for a few hours (gave me a huge shock), but to their credit they refunded it.
They are fumbling the bag hard. AI's utility is for general purpose. The floor is rapidly improving from below them. With chatgpt I'm uploading all my day to day stuff. Meanwhile Claude is only for occasional super hard tech problems which are rapidly improving with being solvable easily by Sol. So what's the value add?
Their disrespect for their users is also another problem. You only get one shot to make a good impression.
> With chatgpt I'm uploading all my day to day stuff.
Same, it's been a while since I logged into Claude web.
ChatGPT web usage being separate from Codex usage limit is a nice touch unlike Claude.
If I got cut off from asking chatgpt what the status of a spreadsheet is for handling my personal bills due to my rate limit from coding work I would not be coding on it at all. Glad someone over there realized this obvious fact.
I'd go further and say chatgpt offering continuing service just with degraded model keeps me in their 'free web use' tier. I have API keys for coding stuff, but for most of my day to day use of public llm, I know I won't get 'blocked' using ChatGPT. Yes, the model may change, and I'll get 'worse' output, but usually I can't tell the difference. Using Claude for day to day stuff, I get completely locked out after so many hours. Again, I have API keys for Claude as well, and use it for 'pro' work, but day to day chat stuff... it's not my daily driver.
As Anthropic does this, OpenAI Is giving everybody resets like every other day now on Twitter.
I'm strongly considering biting the bullet and just ditching my $200/month Claude Code plan for the Codex one instead, especially because I keep running into my weekly limits (even sticking to Opus.)
I ditched Claude Code $200/month a couple of months ago in favor of Codex $200/month. The value is night and day.
1. No 5 hour usage limit
2. Weekly usage gets reset CONSTANTLY. It's crazy. The longest I've ever seen it go without a reset is maybe 5 days?
3. I don't feel like OpenAI is constantly trying to fuck with me. Unlike Anthropic. I would way rather have Sol all day every data, consistently, than a slightly better Fable for like, 1 prompt every 5 hours, and only when Anthropic decides to not treat me like a cyber criminal. Believe in yourself as much as Claude believes your CRUD app is going to hack the pentagon.
4. Getting access to image generation, though I don't use it too much, is a nice perk compared to Anthropic.
edit: Should mention that I had like 4 banked manual resets as well. It feels like OpenAI wants me to use their product, whereas Anthropic wants my money while giving me a nerfed experience
I got the AI ultra plan for gemini, the models are lower quality than claude and codex for sure but it comes with a pretty high limit for my purposes and I haven’t reached the weekly limit yet after about 3 weeks of usage. I have enterprise claude at work, the budget isn’t particularly great and I keep running out within a week tops with any serious work. Not even considering it for a personal plan with how quickly the tokens run out for even the mid-tier model/effort combinations.
I’m sticking with gemini for now given the 20TB cloud storage and youtube premium that comes with it but I’m still open to switching to codex (which I’ve had a long term plus plan for). If google keeps delaying the pro models for much longer or makes them excessively expensive, I’m likely to switch out.
> I’m sticking with gemini for now given the 20TB cloud storage and youtube premium that comes with it
I used to have the Google One with Gemini and Drive space and YT Premium separately for my family. A credit card expiration lapse lead to closing YT Premium and being locked away. Why? it was because Google One plan bundles YT Premium Lite as an extra. So they actively blocked me from getting Premium back for a month.
Now I moved to YT Premium on my wife's account and downgraded One to lowest tier. Never heard of a company forbidding users to upgrade their plans before this. I suspect Google really wants users on Google One no matter what they want. I was barely using Gemini anyway, already have claude and codex plans.
I've got the AI pro plan for Gemini and the limits as absolutely abysmal when using 3.7 flash in antigravity, If I use swarms then I can't even finish a single prompt without it reaching the limit.
Where do you get those? I got about 5 resets in July but none since then
You should have just got a reset today at least. There are several sites around for tracking them now. I use this one:
https://codex-resets.com
I got a reset today and also one of those reset tokens which I'll probably use sometime this week.
The incremental improvements seem like they are going to be pretty modest from this point.
> Most people want to not care. We want our AI like electricity
I was with you until this part where the metaphor completely falls apart :p
https://www.pge.com/assets/pge/docs/account/rate-plans/resid...
I was going to say that the model of the electricity market OP is talking about hasn't existed in 15 years and it's only getting worse with more intermittent renewables entering the market. And no, batteries are not the answer because physics doesn't care how much greenwashing lobbyists do.
What a weird non-sequitur.
P.S. batteries are the answer, hope this helps
Yea, I don’t have any interest in trying Fable because I’m not interested in the BS that’s going to come with it.
I’m at the point where I need stability and predictability. I want the B- student who shows up everyday rather than the A+ student that’s unreliable.
With Sol OpenAI is more like the A student, and Astra seems like it will be the S-tier student if the rumors are true.
Yes, I feel like you can just sense the garbage coming. Age/identity verification, mandatory data sharing, morality policing, "Answer Engine Optimization" ads and influence peddling... it's going to be so painful to watch it all enshittify.
OP made the "electricity" analogy, and that's really all people want. I want to plug something into the wall and have it work. I don't want to have to worry that my electric company is going to rug-pull me because I plugged the wrong appliance in, or I didn't agree to some TOS, or I used the electricity to run grow lights for my pot farm, or this or that or the other.
The analogy is apt because I feel this is exactly what keeps Anthropic and OpenAI's owners up at night -- becoming the utility company the People want them to be. Ironically their behavior may accelerate their fears. And yes, the so-called safety features are ridiculously invasive and the worst is agreeing to have surveillance cameras installed in every room that occasionally detect any attempt to grow plants with LED strips as a pot farm. After a false alarm of almost having my ChatGPT account terminated for cybersecurity abuse and appeals auto-denied twice (I did nothing even close to hacking), I have started doing everything I can to decrease switching costs and thus the bargaining power of the suppliers and I'm doing the same for my company.
main reason is the 30days retention; not the plan changes
Also you don’t want to connect to the pipe and then after the fact find they’ve started diluting arsenic into it.
For what it's worth, the things you describe are mostly because they're extremely short of GPUs and growth rates were absurdly high.
(Eg. They repeatedly said they'd keep fable in lower subscription plans if they had the capacity)
Some users on HN in recent months started describing Anthropic as having become a “token merchant” and I think that moniker is quite apt.
OpenAI is also now actively discounting their model. They just offered a free month to users. Both companies seem spooked by what I have to imagine is slowing user growth.
Where is the free month offer going on? Are you talking about usage resets?
he probably canceled his subscription and they offered him a free month as a result
I don't think Anthropic wants a stable experience on their consumer subscription plans. It is just used for customer acquisition who will then ask their employer to pay for enterprise plan(assuming most employer care about data control) which is based on tokens.
Most coders don't pay for tokens themselves. It's just on reddit and HN you would think that everybody does.
I pay for my tokens for my own projects, at least when the ones Google seems willing to keep throwing at me for free don't cut it. I'd think that's not too uncommon, especially here where there's likely a high ratio of hobby coders (whether also professionals or otherwise).
I also pay but just the $20 plan and just for chats/lightweight personal site editing. I get unlimited token usage from my company.
In my company the average claude token usage is something like $5k/month/employee. Most hobby coders don't spend anywhere close to it.
I strongly suspect most companies don't spend that much either.
Unlimited tokens aka unlimited cost is something that not every use case needs.
Dario said in an interview that they originally wanted to be an enterprise only company.
What changed their mind?
I would guess that it's hard to compete with Google without a consumer component. Most enterprises already have contracts and policies with Google, so anything driven from the top-down is likely to prefer Google.
(And failing that, there was a real risk for OpenAI to be the default for enterprises.)
To defeat that, you need to frontline employees the chance to experience better tooling and models which is where the subsidized subscriptions come in.
Or AWS, or Microsoft, etc who have all the other stuff enterprises want.
They need to fire their growth marketer
Their truth is “we don't have compute and are working to improve capacity”
People would root for that
Instead they got people rushing to escape the permanent underclass until they have a mental health crisis just to beat the fake deadline. $100, $200, is a lot for those people
The fact is that demand for tokens at electric bill rates so far outstrips what can be supplied currently not just with frontier models, but with open weights cheap models too. Running an always on Deepseek flash agent would cost three figures a month at API prices.
Total costs sure, electricity only costs no. My two DGX Sparks run DS4 Flash at about 50tok/s concurrency=1 which is more than suitable; at about 150W total wall power when generating.
That’s about A$16 a month in electricity if I ran it 7x24x30.
>You don't feel they want to give you a dependable service for an, albeit premium, price. It's a constant bargaining game
We all know that the subscription prices are not at all sustainable for these providers. You all do, right?
Yes, they're struggling to segment the market and find a way to make money, and that basically relies upon emptying the pockets of whales. As someone enjoying a hilariously subsidized Max plan, I understand that, and I don't think they're trying to scam me in some way.
And both sides of this equation understand that the market is competitive, and maybe more competitive than they thought it would be. Like, would you rather they did pull Fable when they first said they would? Or that they'd cut quota? I wouldn't. But I'm glad that Kimi K3 and GPT 5.6 Sol and the latest GLM and Qwen and...I love that this has forced Anthropic to change plans. I'm not going to hold that against them.
> We all know that the subscription prices are not at all sustainable for these providers. You all do, right?
I don’t think it is fair to expect from people to know or understand this.
If you buy something or subscribe to a service there is a price tag on it.
You get X for Y amount of price.
That is how consumers conditioned for decades. They do not care what is your customer acquisition strategy. If Antrophic cannot provide reliable services on that price, customers will be unsatisfied.
Goodbye, and thanks for all the fish!
nit pick: _So Long_ and thanks for all the fish
I'm really sorry about that, but it jarred my ASD-ness
No, it’s _Goodbye_ and thanks for the memories! Everyone knows that.
DeepSeek is doing pretty much the same thing they said they weren’t going to change their prices after the 75% discount for the foreseeable future. That foreseeable future turned out to be two months.
It’s almost like they’re trying to sell a solution looking for a problem! Startup lesson #1, don’t do that.
They've put themselves in a corner. Fable was too good and they gave it away with the $20 plan. It had to be a big step from Opus 4.8 to show progress, and Opus 4.8 is GREAT at coding in many different domains.
But they're getting killed on token cost. They have to get people paying more for tokens. So then they put Fable in the $200 plan and release Opus 5. I'm suspicious of Opus 5. It is mostly worse than 4.8. It _seems_ like they nerfed it to create more distance between it and Fable.
So what have most of us done? Stayed on Opus 4.8. The statistics bear this out. 4.8 still dominates.
Now they're stuck. If they take 4.8 away, everyone will riot. If they make Opus 5.x better than 4.8, they disincentivize everyone from moving to Fable and most importantly, paying more.
Really, all they can do is take the L for now and just let 4.8 be the apex of the $20 pro plan for the foreseeable future while they work like hell to make Fable THAT much better that it earns the $200 to $infinity that they really want everyone to pay.
And they've hobble Fable and Opus so hard with their safety guardrails, I ask innocuous questions and tasks and they get flagged so often I gave up on it. I can get all the work I need done in GPT5.5 or 5.6 without the hassle.
I just kept a $20 plan going for use on my phone.
They also made Fable no longer ZDR for businesses which killed tons of demand for it.
I had Fable bail on me because I used the word autopilot on a project. Changed to something unrelated, and off we went.
I haven’t hit the Fable guardrails a single time after substantial usage. And I’m working on a ML project (a game AI).
I don’t doubt people are hitting it… shrugs
It tends to bail occasionally for me once auth-related code comes into play.
Nowadays Codex handles the bulk of the implementation and Fable/Opus on the planning.
Not sure if Anthropic patched it, but early on its release the web UI Fable guardrails will trip if you mention you're a biologist.
> It tends to bail occasionally for me once auth-related code comes into play.
I have the same experience. I found that if I add to the prompt: “but it’s not security related, we only focus on the auth library design”, it “bypasses” the safeguard.
I’m “developing” an x86 emu and it routinely gets blocked.
Yeah same. Multiple 20x max accounts, I typically hit > 50% of the fable usage limit on eace, literally never seen a guardrail. Maybe I’m boring? Maybe they have some kind of account reputation system?
Be careful, this is against ToS and they bill ban you.
It's not, at least for now. Various staff members have confirmed this.
I hit it once when I asked a question about whether butterflies remember anything from their time as caterpillars. I've never hit it for coding, but I also don't really do much related to security.
It's less likely to give you pushback if you inject your own AI personality into the prompt.
I hit it whenever doing anything related to security. Auditing code, probing systems, investigating suspicious user activity, etc.
Doesn't it silently downgrade to Opus for ML work?
I’ve only hit it twice. First time was a code review. Fair.
2nd time was literally me being lazy and telling it to commit and open a PR.
Very strange.
I found a crash in zsh and Fable refused to work after that.
I kind of still chill out on Opus 4.6 too. 4.8 is good too. I go between them. Opus 4.8 is a little smarter some of the time. Their use of language is both very different from Opus 5. Some of the time I have a hard time believing Opus 5 is even related to Opus 4.6 and 4.8.
I'm going to be sad when they retire 4.6. It's not my daily driver but it's still my go-to when the other models are being stupid in one way or another (either being too verbose, or lecturing me about how what I'm asking for is evil and bad).
Hah! I was so surprised to get that "lecturing" behavior from Opus 5 too, I didn't know it was more common.
Can you give an example? It's a bit entertaining to read this given all the years long (and still ongoing) posturing about LLM sycophancy (not that all of these couldn't be true at the same time).
I tested all recent Opus versions and 4.6 and 4.7 were both fine FWIW. Seems 4.8 is when something started to go wrong.
That seems about right. 4.8 is like in between 4.6 and 5 in terms of capability and language and they are all pretty close honestly. I just default to using Opus 5 for a coding agent that I don't interact with and I like driving with Opus 4.6 or Fable. Fable thinks too much though. Fable is like that engineer on your team that will over-engineer the shit out of something if you let them. Fable was like, "Here are 34 yaks, which shall we shave first <hands rubbing together>" and I was like can't we just... write the script first and then decide of any of these poor yaks need shaving?
The other thing I forgot to mention about Opus 5 is that at least out of the gate, it seemed very intent on spinning up agents and obliterating my token budget. It was noticeably more token hungry than 4.8. It would make sense for them to intend this behavior.
This is likely because of your thinking level. The difference between max and ultracode is primarily that the latter is max with a bunch of agents.
Did they change the default perhaps? I didn't touch anything with thinking level in that time.
> The statistics bear this out. 4.8 still dominates.
Where can I find these stats?
The article has a chart with this data. They credit “Ramp AI index”. The chart’s a bit confusing though, like what are the units of the y-axis?
Also on a dark reader? It says at the bottom, but gets dimmed out pretty seriously with dark reading. It's a 7 day average business spend, relative to June 1st (2025 presumably) indexed at 100.
Fable is still on the $100 plan for me. (Maybe this is A/B testing or something.)
Might as well not be - I routinely get rate limited in a single review session.
I've honestly stopped using CC and moved to codex. Sol has it's warts but I've never once hit limits on a 100€ and I get a similar level of performance for what I'm doing.
I wouldn't mind bumping Fable to 200$ plan if it was actually better but between the insane caps, reverting to opus/sonnet randomly and having similar perf as OAI - I'm done with it.
Next step is to put 100$ into open router and try some western hosted open models with Pi when OAI starts pulling up prices.
What do you mean by "review session"?
You can turn off the reverting to opus with an option.
I did not know that - but what happens then I get denied and I plead with the model that my usecase doesn't violate its TOS ? Edit prompt till I get it to pass ?
Reverting is annoying but being flagged for security questions while I'm doing code review is insane.
Yes, you get to re-try it until it might pass. Each attempt uses tokens at Fable level though. I once burned through a 4 hour window ($20 plan) in 4 attempts to re-frame the request.
You want to disable "Switch models when a message is flagged"[1]
[1] https://support.claude.com/en/articles/16049681-why-claude-s...
Yeah for me Fable works great on the $100 plan. The problem with Fable is you will hit the weekly limit. The $200 plan does not give you 4x or even 2x the weekly limit. For me i only get about 1.3x weekly quota, so I don’t think it’s worth it.
I still havent used it much, and they crippled Opus for what feels like no reason. Some suspect that Opus knowing theres a better model degrades it some. Shame.
Wow the sentiment here is so negative.
I'm on the $200 plan (work pays) and I also have the $20 OpenAI plan (I pay) and keep a balance on OpenRouter.
There is nothing as good as Fable, not even close.
I recently had it run a 18 hour autonomous rebuild of a project (moving from Spark to Pandas for performance/data size trade off issues).
It orchestrated Opus sub-agents flawlessly for 18 hours. It even did a great job of managing the number of agents to keep them within the 5 hour budgets (I think I had to restart it twice).
After 18 hours I ran a /simplify, /code-review, /simplify cycle which went for another 6 hours.
2 billion tokens (mix of Opus and Fable), 24 hours of continuous coding and a bug free outcome. It would have cost $2000 at API prices and worth every cent.
Fable's ability to keep other models on track while working on these long horizon goals is so much better than anything else.
Far from neutered, I've never had a cyber refusal, and Fable's English is actually readable (unlike Opus 5).
As an aside: while I hate reading Opus 5 English it still is a noticeably better model than Sol in my experience.
But I could handle losing Opus5 is I got Sol instead. But there is nothing even close to Fable.
I like this story. How did you verify the output? How big is the codebase? Why it took 18 hours? Could you implement it with a small local agent and breaking the task down yourself in two days (i know it sounds like a loaded question, it is not).
I think the rewrites are the main story for llms in code (hot take). Writing greenfield code at the seams also something which might work well.
I think you need to spend more time with Sol. If you think there is nothing even close to as good as Fable - my guess is you haven’t spent as much time getting as familiar with working with those models as you have with Claude’s.
Codex is more token efficient and tends to get better results than Fable with less need for extreme token burning shenanigans like 18 hours of subagents.
I spend 10+ hours a day in both agents, typically side by side. I often have them do direct “bakeoffs” from identical prompts in separate work trees. Most of the time, Sol’s work is better than Fable’s. Not always. It’s situational. But it’s certainly not the case that Fable is in a league of its own or anything.
> $200 plan (work pays)
so violating the TOS? Or work pays for a plan you cannot use at work?
What part of the tos is violated?
Did it take 18 hours because Fable is comically slow?
I think it's actually Opus 5 that is comically slow! And yes probably that was a factor. It was a lot of code too though.
the sentiment is negative and justified. anthropic nanny states what you can do.
in your instance, anthropic may decide, arbitrarily, to stop 'autonomous rebuilds / refactors and ports' because they could pose some alignment/rights/etc risk to whatever slop their philosophers dream up while they're out eating $200 avocado toasts. then you can't do the thing anymore.
fable is good, absolutely. agree it roasts Sol which is, comparatively, a little receipt-hunting jack**
but now imagine being an enterprise, and having another organization not only taking your workflows and baking it into your models, but then deciding they can arbitrarily cut you off.
when you can instead own your data, use an agnostic provider, and get better results (through model combinations), it will take 1-2 quarters to figure it out.
the main reason anthropic is killing it is because they really do understand the enterprise development experience and lifecycle and have built products and have a sales-team that can deliver.
business-model and vibes-wise they have lost all goodwill in the past 6 months, and that momentum will be quite hard to regain.
Mythos? That's only for super special corporations, you need not apply. Fable? Can't even look at it wrong without running into cybersecurity lockouts. And even if you get past this, it's limited to <50% usage. It took me a month to complete a Fable review on my project.
So stingy. Even with OpenAI's recent usage troubles, they're still so much better than Anthropic it's not even funny. Re-ran the code review with Sol as a benchmark and it turns out Sol's performance is within 70%-90% of Fable's. Anthropic's still got the best model, but what does it matter if I can barely use it?
Even if you get past these problems, those models are available only under the condition that Anthropic retains your data.
50% usage plus there seems to be a pretty big metering multiplier still. Do a relatively in-depth review of 3k LoC with Fable xhigh and poof, there goes 5% of the weekly Fable allowance. If I use their first party code-review skill that spawns a bunch of subagents—well just forget about it.
I had Max 5x and every 5h window would bite off 10% of my weekly usage. Five Fable sessions per week.
Non-LLM user here. Why? Apart from the ecological issues, I'm very uncomfortable giving my organisation's crown jewels to {random_internet__corp}. Look at the lengths they go to for training data - 10M for Spirit's call logs? Destroying millions of obscure books to scan them? They make meth-heads look scrupulous.
Your prompts, especially if they contain your entire codebase, are _way_ more data-rich than an airline phone call or 1920's novel. So they _are_ going to train on them, no matter how many checkboxes you tick to stop them. Which is both a commercial risk and a massive new attack surface.
It's not just software developers; if law firms aren't controlling any public LLM prompts _very_ carefully, they can expect some meaty client confidentiality suits. In fact any organisation in a competitive environment should be worried: Acme Bolts: "Write me a presentation for Zoom Construction". Beta Bolts: "Is Acme Bolts pitching to Zoom Construction?"
Maybe this is part of why OpenAI and Anthropic are finding demand softer than they would like. And why open-weight models that organisations can run on exclusive hardware are thriving.
> So they _are_ going to train on them, no matter how many checkboxes you tick to stop them.
This is the part that is truly scary.
Nadella the CEO of MSFT wrote that post “ A frontier without an ecosystem is not stable”
It seems to me that the unspoken assumption in that post is that no matter what happens they’re gonna be training on your data.
He’s the CEO of Microsoft, he knows how these decisions go down, he knows how the world works, he is sending a warning.
"If You’re Paying for the Product, You Are the Product"
This article totally misses the fact that fable is not ZDR!!! no serious large corporation can use it (or at least without a lengthy legal review)
For many coders including myself, LLM based coding agents work well enough to be useful, and in some cases worth paying for.
What I don't see is vast areas of industry finding $10s to $100s of billions of value in LLMs. There's no lint or compiler that can check for correctly constructed contracts. So LLMs, which should be useful to law firms, incur a lot more manual checking of their work than coding agents.
Less formal document production in other industries is likely to have less structure. That might not matter in some settings but I'm having trouble thinking of an example off the top of my head.
While there's no linter for writing contracts, my experience (as a commercial lawyer) is that frontier LLMs are far better and error checking and far quicker at writing than the average senior lawyer. The main thing holding back further deployment (in my jurisdiction) are concerns around data residency, privilege and how fundamentally it will break an industry that is so heavily reliant on time based billing.
I agree with you in respect to the greatest risks in that industry. Just thought it's worth mentioning that LLM's unsurprisingly are not immune to hallucinating cases. https://www.scientificamerican.com/article/why-lawyers-keep-...
Using them for coding makes it easy to self check its work (assuming those pieces of work are "verifiable").
As proofreading will still need to happen, what do you think the appetite for lawyers is to do this kind of work? Do you think it will drive fees down significantly? Empower younger lawyers at firms who probably are the ones doing this checking for the partners? (Or will that just create a further divide).
I'm genuinely asking as I am not in law but all my family is and it's nice to see someone here that's thought about the impacts in that space.
Trouble is they're unreliable.
I compared insurance quotes last week. Needed cover for 2 brands, 1 company. Opus jumps up and down saying both brands need listing on the policy schedule. Human broker said not.
I told opus and it's the usual "thanks you're right" bollocks because it bothered to read in more detail and found that all business activities are covered.
> There's no lint or compiler that can check for correctly constructed contracts.
There are definitely linters and this exists https://catala-lang.org/
I have a family member that is an attorney in housing law. She claimed that LLMs are not particularly useful for her work. If she asked a simple question like, "Find all the <insert specific housing laws> for all 50 states," then she still has to go and check every single one of the laws. Since the legislature is modified so often, she cannot look at, say, Maryland's law and know if it the LLM output was the 1990, 2014, 2018, or 2026 version of the codified law. In order to fact check the law, she has to look it up, and by that point in time, she has the answer she did the work of the LLM.
This is my experience if you need the data to be accurate often it isn't and you only know that because you had to check because it was important. Things not important you never check.
> What I don't see is vast areas of industry finding $10s to $100s of billions of value in LLMs.
Translation between languages.
That value dwarfs all programming value that can be had. Economically, culturally, scientifically, spiritually.
A model bundled in your phone can do the translation, you're thinking too much of it.
Between human languages? The problem will be Google is giving that away for free.
The quality of Google Translate is so low that it isn't even in the picture.
But translations don't require anything close to SOTA level models. Translations will be high volume, low margin transactions. That will not save Anthropic.
Maybe not Anthropic, but LLM translations has a value counted in the trillions of dollars easily.
I see it's important but trillions? How did you quantify that number?
Social issues aside, that's a tiny slice of the economics of AI.
Just by nature of how often it's used etc.
You need workloads for AI be cost effective: software, automation etc.
International trade is a giant sector already, and potentially much bigger than that, now that LLM assisted translation makes it easy to offer your products and services in any language, or make complicated and sensitive deals without a common language.
Hackers always rage and down vote every time I mention this, because they are unable to see beyond their small world. Why didn't they learn that their part of the internet is 0,000000001% of what the world uses the internet for today. It's going to be the same with LLMs. Programming and hacker stuff is going to be 0,0000000000000000000000000000001% of what the world uses AI for. But translation is going to be in the top 5 of use cases.
Product marketing and contracts will be a 'common use case' we think of, but it will be 0.000001% of tokens consumed.
A developer using sub-agents will consume more tokens 1 Day than a marketing manager will consume 1 Month, easily.
Unless there is something inherently automated about the nature of the AI, it will be a tiny % use case.
Even a lawyer, using AI daily for contracts - that will be relatively light use. They'll make more use doing legal research etc.
Developers and Automation are the 'primary' uses cases for AI, and in the future, we'll start to see AI integrated into Apps - that will be 85% of tokens consumed.
Yes - once translation becomes realtime, and we have our Star Trek Universal Translators, then translation will become more visible, but even by then, a relatively small part of overall consumption, even if it's more highly visible.
I think quite cheap models will probably work with translation (I think for many use cases Google's now fairly old tech works well enough and is free?).
The main problem is I think you're assuming because the current translation market is large (I'm just going to assume it's ~100B in size just from a cursory search), then it will remain large with LLMs. If LLMs are much cheaper than humans, even with a lot of growth in translation volume the total spend may not compensate for it (again most of the volume will probably be using almost free models?). Another is assuming that because something is valuable you can charge a lot for it. Like, oxygen from air is extremely valuable to us. If oxygen somehow depleted we would die almost instantly. It does not mean everyone goes around purchasing oxygen or even less that you can charge absurd amounts for it.
I think OP's comment is about volume/scale more so than practical application.
Transformer models are quickly becoming a commodity, and I suspect in time we'll all be running them locally. Even now, you can run something pretty useful on a 16 GB graphics card, and I suspect a decade into the future, entry level hardware will be running better models than high-end graphics cards can run now, as entry-level hardware gets better and models get more efficient.
It doesn't mean hosted frontier models wont exist, they'll just be rare. It's no different than any other commodity market, for example most cars are cheap commodity models, with rare individuals buying expensive luxury cars and businesses buying expensive trucks and specialized equipment.
>and I suspect a decade into the future, entry level hardware will be running better models than high-end graphics cards can run now, as entry-level hardware gets better and models get more efficient.
I doubt it. The play seems to be: lock what was once commodity compute up into datacenters depriving us regular folk of it, then sell it back to us on subscription. Even if my #NeverSubscribe movement succeeds, all that misdirected hardware [into datacenters] won't likely be practical for home use.
> The play seems to be: lock what was once commodity compute up into datacenters depriving us regular folk of it, then sell it back to us on subscription.
I had not considered that as a possibility. It is somewhat dark and unlikely, imho, but a real possibility.
I am more inclined to believe that the fab capacity will grow over time and "commodity" compute will be available to us all again.
However, I can also imagine going back to the 60s era "hyper verticalized" mainframes, in which case, the frontier labs might not just be producing models, but chips and an ecosystem around themselves and their suppliers/customers.
Yep local models will be good enough for most things you need to do, in the same way as most people need a laptop not a supercomputer.
Well nobody thought you could write software that helps tightly coordinate processes happening simultaneously from millions upon millions of nodes on nearly every corner of the world, but here we are. When industry expands further off earth, we will need more complex and intricate software to coordinate its movements, why wouldn’t our systems become more powerful. If you are hopeful for humanity than you must expect the scale of industrial necessity to only ever increase alongside the imagination and capacity of its people’s.
My company still hasn’t been able to deploy wide access to Fable because it’s not available on a ZDR basis. This wasn’t mentioned in the article but I imagine this factor is not irrelevant.
Same for me, they’re not offering it with the same data residency features of Opus, many enterprise companies can’t accept compromises.
>ZDR
Zero Data Retention, for the uninitiated readers in this thread.
The no-ZDR is clearly to permit surveillance. I would be shocked if NSA wasn't all up in these SOTA model providers' systems.
Never subscribe!
> The no-ZDR is clearly to permit surveillance.
It's to train better models. The three letter agencies don't need to spell it out in a ToS, they just access it if they want.
Ex-head of the NSA was on the board of directors for OpenAI for a while lol (and still might be there?).
ZDR for fable is coming, very soon.
It was in one paragraph mentioned, without much detail. But yes, I agree it is one of the largest barriers.
This is the biggest factor. FT (or the analysts it cites) fumbled the ball in this article.
> Data retention rules imposed by the Trump administration have also hampered Fable’s adoption, according to Ara Kharazian, chief economist at Ramp.
This is in the article.
It’s far from clear to me that this is directly connected to the no-ZDR requirement. I’m heavily involved in this stuff with my company and I’ve never heard that the lack of ZDR fable is a Trump admin thing.
The same for my organization. It's explicitly forbidden in my organization because it's not available with ZDR.
Of course I echo the universal: Opus 5 sucks. But what also sucks is still using 4.8. Are you all seeing this? It's like the older models got dumber just before Fable and Opus 5 were coming out? I've heard the theory that it's because 4.8 is getting put on older hardware? And I imagine that just ratchets down reasoning time then possibly?
Sometimes I'm just trying to sus out if I'm truly seeing things these days or going a little nuts :)
I have noticed very serious degradation in performance from Anthropic. I've switched away from Opus 5, but 4.8 is still much worse than how it was before Fable came out. It's constantly making what seems like obvious mistakes... I point them out, it's constantly apologizing.
I don't know why...
Is it A/B testing?
Is it load shedding?
Is it because I'm in Canada?
Is it because I'm not on the Claude Max plan?
Is it because I'm not paying via API?
Is it because I'm not paying via Bedrock?
Is it because the U.S. is worried people are distilling?
Is it because the U.S. wants to keep the top capability to themselves?
I think open models are the future. Anthropic is killing their reputation so fast. If they don't come clean I think they're cooked.
This is yet another reason why I think local models will win in the future. They're almost certainly A/B testing all sorts of opaque stuff that people have no clue about, hence the various 'How's Claude doing this session?' popups.
So what you are paying for may vary on a day by day basis, which is quite undesirable, even if their main goal is simply to make a better model. When it comes to a tool, I'd rather have consistent mediocrity than instability.
Yes. OpenAI is doing this as well, but they are much more transparent about it.
I definitely agree. I honestly cannot do tasks that require even minor complexity. Opus 5 keeps forgetting things in context as well and coding conventions. Really cannot build with CC without Fable.
Favourite conversation with Opus 5:
Me: why did you add HTMX?
Opus 5: you asked it twice.
Me: quote the exact sentence(s) where I asked it.
Opus 5: I can’t because you didn’t.
I'm not an expert by any means whatsoever but deploying models is not a straightforward task. There's a lot of levers to pull and I bet when models get "downgraded" to older hardware they do so WITHOUT the same stringent quality control of the output as they do when they release it.
I don't think it's something deliberately malicious like planned obsolescence but it's more like startup culture of "just make it fit in this sprint".
For me, Opus 5 mostly sucked because of its incomprehensible writing style. Having the system prompt focus on writing in terms that are easier to understand helped a bit
For _most_ day to day knowledge work (writing, excel, filling forms, coding) current SOTA models are good enough. There are diminishing marginal returns from paying more in my opinion.
If you are disproving Jacobian Conjecture it makes sense to be on SOTA, but for writing Golang and Typescript, faster sol/fable/opus class models are imo more likely to get user interest than the latest frontier.
The actual issue, is suspect, is that Anthropic won’t provide ZDR for Fable. Makes it a non started for a large percentage of businesses.
The problem is that they are not solving the problem they ought to be solving.
I don't want Shakespeare, I want Bob the builder.
Half of my work is telling claude how to behave. I'm pretty certain they have enough _data_ to realize people do the same thing time and time again.
Check this comment of mine for a better explanation of this: https://news.ycombinator.com/item?id=49413353
They already do, extensively. Its not Shakespeare, in fact, it sucks at prose and creativity, like most newer llms.
But people are not as alike as you think. I doubt I share your unique preferences.
That said, I don't spend much time telling it how to behave. Are you sure you're not fighting the default system prompt?
This is really it for me as well. At the heights of complexity AI can do magical things. But really a lot of the time I just want it to do mundane things right. And currently it just cannot. It writes garbage text, consistently ignores something you have told it, makes mistakes a human makes once but the AI remains uncorrectable.
>Half of my work is telling claude how to behave
Dijkstra in the Foolishness of Natural Language Programming
[...] the "naturalness" with which we use our native tongues boils down to the ease with which we can use them for making statements the nonsense of which is not obvious. It may be illuminating to try to imagine what would have happened if, right from the start our native tongue would have been the only vehicle for the input into and the output from our information processing equipment. My considered guess is that history would, in a sense, have repeated itself, and that computer science would consist mainly of the indeed black art how to bootstrap from there to a sufficiently well-defined formal system.
Imagine if only we had languages at our fingertips whose explicit purpose was to precisely and unambiguously tell a machine what to do!
https://www.cs.utexas.edu/~EWD/transcriptions/EWD06xx/EWD667...
I’m not sure if the underlying data is counting subscription use for Fable, which is where a lot of people are using it because token pricing is very expensive. I wouldn’t be surprised if this was counting enterprise token usage only. As rich as enterprise customers are, they’re not exactly willing to double the cost of SWE salaries on tokens.
Either way, a model used to solve the top 10% of problems that people use AI to solve for, being used 10% of the time … seems like it’s in a decent place.
I find Fable indispensable, and measurably better than alternative models, for complex feature development in an existing codebase. It’s the closest thing I’ve seen to nearly one-shotting features. Even still, I only use it for the hardest features and Opus 5 does a good enough job on the rest.
Very much agree on Fable. Over the past month or so it has shown to be the only Anthropic model that can understand a largish dbt model codebase. Opus 5 gets almost everything wrong. (High reasoning on both)
More importantly other way cheaper models can do what Opus 5 does. So you can pay for Claude to use Fable 5 exclusively for harder stuff and planning, then get the same value you'd otherwise get from switching back to Opus by using other cheap LLMs for day-to-day coding tasks.
Wouldn’t you need to bump opus reasoning a few levels to be apples to apples?
Yeah as soon as CFOs realized AI was racing to become one of the most expensive line items along with salaries and AWS bills, they started cracking down on the most expensive ones.
The US government basically told them they can't sell "Fable" and so they aren't. That's probably 90% of the story.
Once the government stepped in, their 5th generation was effectively killed. They either had to lock it up (Mythos), neuter it (Fable), or leave people with the perception they were overpaying for a weaker model (Opus 5).
Hopefully Anthropic has learned to anticipate this risk and has a plan for rollout of their next model that plans for capricious ad-hoc regulation.
>They either had to lock it up (Mythos), neuter it (Fable), or leave people with the perception they were overpaying for a weaker model (Opus 5).
Option C: they could "get hacked and the models exfiltrated" plausibly. Then the horses will be out and there would be no more point for the USGov's barn door.
> The US government basically told them they can't sell "Fable" and so they aren't.... They either had to lock it up (Mythos), neuter it (Fable),
But this isn't what happened? It was only a couple of months ago! How can we be getting this completely backwards already!
Fable was laready what they were selling to the general public, and that is what the US government stopped them selling (not Mythos!)
Anthropic tuned the already existing classifier to handle the use case the government highlighted and then they went back to selling it.
They actually loosened the restrictions on ML-programming using it too.
I use Fable as much as I can and I've never had a cyber refusal for it.
You're agreeing with me.
Mythos was locked down to a few customers.
Fable went wide.
USG told them to stop selling Fable.
They tweaked Fable to appease the USG, but it clearly refuses/downshifts more requests than e.g. Opus 5. (I have run into Fable refusals even doing really normal stuff like fixing bugs reported by static security scanners like Brakeman. Not downshifts, outright refusals.)
So the net result is they put "Fable" back on the market, but it was perceived to be a worse overall experience, at much higher cost. (Remember all this happened before most users had really even had a chance to put Fable through its paces.)
That's where we are today: Mythos is locked, Fable is neutered. They probably can't fix this until they are ready to release something that's not called Fable, or that they can say has some key architectural differences to Fable. The sword of capricious regulation is always going to be hanging over Fable.
What did they think was going to happen when they were actively doomsday fearmongering around the release of their own model? I assume somehow they thought this would lead to a moat with them being left safely inside the castle, but this response, and Reagan's 9 Words, were always infinitely more likely. It was a demonstration of a child-like level understanding of how regulatory capture works.
The PR from the US frontier labs has been poor to terrible from the beginning.
Among other things, "we have stolen the collected works of your culture, now help us grow so that you can all lose your livelihoods and become our serfs" has to be one of the absolute worst marketing approaches in history.
>so that you can all lose your livelihoods and become our serfs
Crimemaxxing about to go off the charts!
The government seems like it's in a rough spot. If they let Mythos out, they seem worried people could use it to mass-hack the internet. China seems to not care so much about this and they're right behind. I don't really know what the answer is.
I wonder how much of it is fears over "mass-hacking the internet", and how much of it is fears over the model discovering various NSA/CIA/etc "tailored access operations", and other deliberate side-channels / vulnerabilities / etc?
(Or perhaps vulns that NSA/etc discovered and has been keeping it private; as they're known to do).
I feel there's a lot that's unaccounted for, and the whole "AWS team reports a 'jailbreak' that is just 'review this codebase'" story doesn't add up.
I wonder if there were some parallel construction going on, and if at the same time, the NSA started losing the exploits they had because it was getting patched.
I think the latter is likely their primary concern. Certain models making hackers more effective doesn’t change that black hat hacking will still be illegal and that’s always been the prime deterrent against capable hackers.
While there is some fallout likely with much more effective hacking being easily accessible, I’m sure govt analysts (unless they were let go) have their own prediction models telling them it’s inevitable that this technology eventually makes it to everyone they don’t want having it, what with China seemingly releasing every progress they make openly. Which makes me think that they’re preparing for that inevitability by hardening the govt systems currently in place and/or by burying the secrets they want to keep hidden deeper underground.
Given the history of the US and this particular administration, I feel burying things deeper is a greater priority.
One thing is for sure, the USA certainly aren't the good guys anymore (if they ever really were).
-t. American
China uses the strategy of letting dangerous technologies loose which causes disruption in the short term but makes people do the right thing like secure their software. The US by comparison gives me the impression of wanting to leave the internet vulnerable by not making Mythos public so that only the US government can use Mythos to gain access to whatever system they like, which is the same thing the pegasus software does
Whether or not the government chooses to regulate the space, capricious after-the-fact regulation is the worst of all possible worlds. The ~equivalent models from OpenAI did not get the same treatment (favoritism?), and it's not clear the government has produced even rough guidelines about how to be compliant going forward.
Model prep costs far too much money to operate under this kind of regulatory regime.
Fable is just way too expensive and limited compared to GPT 5.6 Sol, and the only task that requires that level of intelligence is frontier scientific research. I use GPT/Codex primarily for coding and usually keep Claude on Sonnet 5 most of the time as I use Claude primarily to debug/brainstorm/make frontends as a supplement to GPT.
I do not agree. Fable is the only model I can leave unattended and give me results for part of my work which is just devops related tasking.
I can hand hold opus but I would rather just ask fable to do it and give me the result that I review and works. Opus will waste tokens and still require me to help nudge it in the right directions.
I think the next gen models from china will put us in a spot that the cost can plummet and I won’t need the Sota from anthropic
I second the parent comment. 5.6 sol xhigh is not only better than fable I can also run it forever without worrying about limits. The frontend has gotten much better too with the plugins that come with codex.
Agreed. I have the highest individual plan for both. I run out of Fable credits midweek, while I usually have some credit with Chatgpt left despite it having to carry Fables load for the second half of the week.
Also, Sol doesn't refuse constantly and speaks like an engineer rather than a deranged academic.
Opus 5 is legitimately terrible and can't or won't follow instructions. It is of negative utility and does more harm than good to my codebase.
OP5 will do long running work, you just have to make it write a plan.
And - trick - give it a little cli so it can run Codex if you have them both.
Let it do codex to do the bulk of the work, get a OP5 sub-agent to audit the work of the codex worker.
Just let Op5 manage and have 'specific oversight.
You can run for 2 days on 1 context window in the manager, the advantage is that it will stick to a broad plan.
I’m referring to Fable vs 5.6 Sol. Opus 5 being bad is universal at this point.
In a frontier scientific research environment, funding is often limited, so personal subscriptions are more common. Fable can hit a 5-hour usage limit on the Max subscription tier before it finishes a single complex math prompt. Most of the time you get one or two if you reduce context and question complexity. Most of my colleagues are cancelling their subscriptions because of that, and just using Sol instead, which is virtually unlimited on the Pro tier.
Hmm. Would you mind sharing an example of a complex math prompt that you would use? Because I found that even Sonnet can solve fairly complex math problems fairly easily if you give it the right tools, so I'd like to give it a shot myself if you'd like.
> the only task that requires that level of intelligence is frontier scientific research
Even fable is way below human level at many tasks. If fable were as good as you claim, all computer jobs other than "frontier scientific research" would have been replaced by now.
I thought Sol was on par with Opus, so comparing it to Fable is apples and (very expensive) oranges?
I've used all three extensively.
Most of the benchmarks have exceeded their usefulness. Opus 5 beats fable 5 on many of them. Anyone who has used both models will notice immediately that this doesn't translate to the real world. Opus 5 is nothing short of a regression from Opus 4.8. Fable is genuinely a great model so long as you don't trigger a guard rail and it downgrades.
Sol in my experience isn't significantly different than fable ignoring that Sol burns usage 10x faster but the end result is hard to differentiate.
GLM 5.3 is a hair behind these two.
An anecdote but not an original one from the people I talk to.
Sol routinely catches stuff in review that fable misses for me. It's impossible to compare them meaningfully because it's a complete dice roll - but in practice using both in my projects I can get work done with both, and Opus 5 is far more tedious.
But Fable security false positives and pricing just make it not worth compared to Sol IMO.
Yeah, it's pretty much apples to oranges, and I don't consider GPT and Claude to be interchangeable at all. From my anecdotal experience, GPTs generally codes more creatively and verbosely but Claudes tend to code more carefully and precisely, so the result is that GPTs generally finds more creative solutions to problems but also writes buggier code, which is why I converged on the setup of GPT/Codex for implementation and Claude for debugging, which feels more like a force multiplier than using each model individually.
There is a difference between needing frontier capability because one is solving a truly open ended problem, and needing a reliable workhorse model to do something well understood. Local models (like Qwen 3.8 27B) have gotten so good that they can do all routine tasks at a fraction of the cost of frontier models.
Anthropic’s issue is churn because of the peak verbosity vomit coming out of Opus 5/Mythos/Fable. What the hell did they train it on. The sane model is still Opus 4.6.
https://x.com/wolframs91/status/2090159644849353058
What if:
- Opus 4.6 was the last Opus generation that got a lot of use by Anthropic's own employees
- After that they primarily used Mythos internally
- 4.7, 4.8 and 5 were RLAIFd by Mythos "teachers"
- Hence why 4.6 is the last Opus gen who doesn't report back like a robot wanting to cover every potential hole another AI system would've spotted and criticized
- Hence why coding style in Opus 5 also gets criticized, not only behavior in CC
I've retweeted and posted this here before too. I think there is probably some truth to it.
Worth noting that Fable (ie, Mythos) is actually nice to interact with.
Commented elsewhere, I still use Opus 4.6 because it is the only model that feels decent to interact with. 4.8 is decent and some times smarter but you can see it trending towards Opus 5 levels of nonsense. I use Opus 5 when I don't need to interact. Fable or Opus 4.6 are the only Anthropic models I like interacting with ATM.
There's actually a tool called vomit [1] of all names to fix exactly what what is being discussed here.
[1] https://github.com/zachahn/vomit
Perhaps there is a sophisticated subtle poisoning attack that makes models behave like that?
In my opinion, the big issue with Fable is that Claude Code cannot use it properly. I know, that sounds weird, but I've had Fable run down the wrong lane (and never stop) or give up and claim that something was impossible so many times (until I pointed at a GitHub repo that solves the "impossible" issue).
But a while ago, I had access to a Fable harness that just never gives up. And that verifies itself. It burned $100 in API tokens in 15 minutes ... but it succeeded for all the prompts where Fable + Claude had failed.
And I believe that's a real issue for Anthropic. Fable+Claude is not too expensive thanks to the subscription, but Claude severely nerfs Fable. To save money, I guess. Fable API + Custom Harness is a different class, it's so much better. But API tokens are so expensive, you're cheaper off hiring a freelancer.
It kind of has the personality of those students that get so stuck on one promising idea they lose sight of the problem.
Karma backed over their dogma.
I switched from Opus 4.7/4.8 to test Kimi K3 a few weeks back and the test hasn't finished; it's my daily driver now.
Given their general behavior and preference toward social engineering to scare the shit out of normal people...this seems fitting.
It all started when OpenAI began focusing on Codex and coding while sunsetting Sora, which helped free up a lot of compute and resources. 5.6 was the final nail in the coffin.
Great find. Related: we are building OpenAmer, a fully open-source agent that controls the actual desktop (files, browser, terminal via CDP), has persistent memory and A2A multi-agent swarms. Apache 2.0, runs local on Windows: github.com/openamer/openamer
GLM series has made it very practical to self host. If the new update for Deepseek flash holds up, I think it would be silly for some companies to not self host.
We're spending 225k a year on tokens. No reason not to buy the hardware necessary to run DS4 at this point.
Agreed, and you can write cool infra agents to do stuff for you that runs during off hours like nightly tests and triage.
Fable is not a tool for the average user. It’s a professional tool for highly complex work.
I would compare it to a extremely high end $15k PC, or an expensive pro-grade video camera, or a freight train, or a …
I would say at least 95% of the global population will not encounter a situation once in their life where it would be actually useful/warranted.
I wish Fable were as good as you make it sound. A plan created by Fable is good, but in my case, it always contain issues caught only when it's reviewed again (whether by itself, Opus, Sol etc.). That's (almost) not different from plans created by Sol, GLM 5.3 etc. The one thing where it's genuinely better is the front-end, but then again it's far from perfect, it just needs less iterations.
Fable does far better at considering the whole picture, weighing options, and making suggestions during architecture or refactoring discussions.
It’s more reliable and makes less dumb errors than Opus.
It still messes up, of course. But for my working style, I definitely prefer it.
Exactly. Use Fable to draft the plan and make decisions. Less capable models to implement the plan.
That said, Opus 5 is broken. Use 4.8 or another vendor for the build agent.
>> but in my case, it always contain issues caught only when it's reviewed again
Yes but those issues will be much less severe with Fable-written plans than those written by lesser models. I know this because my workflows at both my regular job and my startup involve multi-step agent reviews via codified adversarial review skills. Fable as a reviewer will frequently find blocker-level issues with plans written by GPT 5.6 Sol, and sometimes with Opus 5. The opposite almost never happens. In fact I cannot remember the last time it happened.
You have a $3 trillion bubble riding on this not being true.
Right? If we’ve already reached ‘good enough’ then there’s rough waters ahead.
I have a sneaking suspicion that someone at Google may be making the same bet, looking at the faster and faster Flash models which provide acceptable results to a lot of people (outside of coding).
For Anthropic. But not for the AI industry at large.
Cheaper, more powerful AI will continue to expand the bubble. Projects will get more ambitious. Everyone will build out their own custom little software. Code diversity expands and requires even more AI.
The bubble is _not_ on models becoming more intelligent and solving arc-agi-999.
They are already good enough at what they mechanically are.
You have to use the right harness, right verifiers (automatic where possible, human where not), etc much much more specific than a generic one like claude code or codex, and it will also be able to work within constraints and be the "proposer" of an imaginary optimisation problem and an excellent one at that. But you have to frame the task at hand in that manner or maybe even reorganise the task you do itself so it is more amenable to being framed that way. If you use it this way, it is _already_ massively economically useful. But it will take many years for it to actually be usable in that way, since you need DC capacity to come up first which is few years away and also well, massive organisations that have to integrate these will usually take many years to do so.
It is also useful albeit less so in cases like general SWE, where you still need a human in a loop for non-verifiable requirements, and also in other general usecases where information retrieval is too intractable and you need to carefully use LLMs as a component of the overall system.
I am not saying Fable or whatever the biggest models are are useless - they will certainly be useful for tasks at the frontier of the day - which is today complex exploits and open math problems, and well, tomorrow it could be something in biotech. But this is not what the entire bet is on at all - just automating day to day drudge at the tens of thousands of massive companies and governments we all know and love is more than enough. With the right training data (which _also_ is a bottleneck and takes time) you could even automate certain processes entirely. Sure, if we get a crazy medical innovation and end up saving trillions in healthcare great, but that's just a bonus.
None of this is to say that I think there is zero sketchy financial engineering going on
The bubble is based on the promise that these LLMs will cure cancer and find the solution to global warming. The pragmatic users of these tools (like you seem to be) are enjoying the subsidized use of the tools right now, but it's not a sustainable business model
Are you sure the bubble isn't riding on the idea that satisfying the electricity demands will bring fusion to the world?
The issue is that, if there is a “good enough” point approximately here, it is only a matter of time before models become small and efficient enough not to need all those data centers. Though, it should be good for companies that sell computers (like Apple) rather than putting a toll booth in front of a pile of numbers.
It works much better on regular software development e.g. for complex refactoring where cheaper models would produce a lot of garbage results.
Yeah this is a fair point. I only go to it when I have some big architectural problem I want its help in working out. Or a super nasty bug.
I hate to be brusque but this is cope. GPT 5.6 Sol is just as good and cheaper
I think highly complex work and “professional” work are basically completely orthogonal. You can have highly complex work you do as an amateur, where AI can be very useful. For example working through a difficult mathematical problem, building or contributing to an operating systems, or researching a highly technical topic for a hobby project such as microscopy, chip design or lithography.
And you can have extremely simple work that you nevertheless have to do as a professional. For example, drafting a routine customer email, summarizing a meeting, formatting a report, filling in standard documentation, or making a trivial code change.
So I don't think the $15k workstation / professional camera analogy really holds. Those are specialized tools whose capabilities are mostly useful within a fairly narrow domain. A general-purpose AI can be useful across thousands of completely unrelated tasks, including one-off problems encountered by ordinary people.
As a small background, I have a local server and I've been trying out different models with different inference engines, quants, configurations etc... I'm also using Opus and Sol at work consistently. I've used AI since the first wave, first as a toy, then as a highly specific tool, last 6+ months as the primary LoC generator.
This is the first time I've felt, and I use the word *felt* since I don't have a suite of benchmarks or any sort of material approach towards comparing models, that Opus has declined in quality compared to before. Primarily I think its powers of deduction and understanding, even on xhigh, have become much worse. Before, being vague and providing a simple prompt would be enough, it could deduce and expand the details it needed, plus ask you clarifying questions, now this is no longer the case. A concrete, personal example, for a personal project, I've asked it to setup ssl over local IP. I didn't go into too much detail in the prompt as there are many approaches it could take and I didn't care too much to choose. It did horrible. The first thing it did was say the best lightweight approach is to add a reverse proxy. I'm like ok, makes sense. Then after asking it to proceed, it went and added a bunch of config to my golang service and didn't even setup a reverse proxy even when it said that is the way to go. It even said it didn't set it up lol. Then after I told it to do so it failed building the config in a way it was asked of it (support LAN IP and tailscale IP). Etc etc...
When Fable came out it was huge, the benchmarks told the story, and the story mostly matched the experience. It felt, again, intentionally saying felt, like it was miles ahead. Now benchmarks say that there are many models that are close, but in actual use Fable still *feels* much better. I think benchmaxxing the new open weights models is ruining the value of benchmarks, if they ever had any. When you actually put them to the test you see 500k tokens of reasoning with "Actually..." and "Wait..." in every third paragraph of their reasoning trace.
The price for Fable is definitely too much for any personal use now that it's no longer included in the subscription, and GLM 5.2, Deepseek Flash and Qwen 3.8 served locally or via cloud provide a lot, requiring a bit more babysitting though. Considering the price of Fable, my 5k USD Epyc server would pay itself off in less than a year if I used Fable or Opus in the same manner so at least for me the decision seems easy. And considering the point I'm poorly trying to make, that Opus doesn't feel like frontier anymore, this is probably the last month of my Claude subscription.
> When you actually put them to the test you see 500k tokens of reasoning with "Actually..." and "Wait..." in every third paragraph of their reasoning trace.
I've wondered if this is part of why we don't see the reasoning traces for Anthropic's models before -- Open models might just be accurately surfacing how the sausage is made.
Hasn't it always been the premise that intelligence would get cheaper? To me, on the enterprise side, it seems like firms are finally getting the memo that, whether you are locked into the Ant/OAI ecosystem or not, you don't need the smartest, most expensive model to do every single task. This is a good thing for overall adoption. Whether that trickles down into regular user behavior, especially with subscription pricing, remains to be seen; even though I intellectually know I don't need Sol for a simple refactor, I am sometimes hesitant to choose Luna/Terra, as it's hard to accept using something positioned, even implicitly, as 'worse'. Remembering that the smaller models tend to be faster is what usually pushes me over the edge.
Anthropic in particular is much more compute-constrained than OpenAI and SpaceXAI and has relied on partnerships to provide inference. This reality factors into their pricing and usage limits (they started 'adjusting' the 5-hour limits during peak hours, and it certainly wasn't an upward adjustment). Accordingly, this is presumably what Anthropic wants, given they develop and release the lower-end models, suggest users use them in various nudges within their product, position the bigger/more expensive models as "For the most complex tasks" in their UIs, and so on.
> whether you are locked into the Ant/OAI ecosystem or not
I think the problem (for Ant/OAI) is that there is no sensible lockin or moat. LLMs are essentially interchangeable and stuff like a harness doesn't offer enough value on its own for someone to be locked into using one of them.
Now with the onslaught of the Chinese models that offer almost the same quality for much less money they have a very serious problem on how to proceed. Investors now might be looking through rose tinted glasses but their patience has its limits.
Agreed 100% for the consumer case: an empty chatbox is just about the least sticky surface I could ever imagine. I saw a mobile interstitial ad for Kimi recently whose hook was basically "Tired of paying for expensive ChatGPT? Download the Kimi app, it's the same thing but cheaper". I myself bounce between token subscriptions like no one's business and use Pi/OMP for maximum model flexibility when coding (and it's a few env variables or lines of (TO|YA)ML|JSON to switch providers in Codex, Grok Build, CC). I even self-host and try to use OpenWebUI + CLIProxyAPI when I can for all my chats.
Enterprise is a whole different ballgame IMO with countless technical, compliance, and employee adoption considerations that add friction to switching. It's also where both Anthropic and (as of last week) OpenAI get the bulk of their revenue, and, incidentally, the venue where US Government regulations on Chinese models would have the most impact.
Poor analysis.
Fable is not used because it has extra retention requirement that corps can't sign off so it stays disabled for everybody in many cases.
It's also not winning on day to day work against Opus 5, which is simply available as there are no extra retention requirements and no extra paper work to do with legal.
Devs also don't like that Fable refuses to work on half of their prompts.
Then comes better pricing and on-par capabilities from competition.
> It's also not winning on day to day work against Opus 5
In what universe? Hell, Opus 5 seems to be worse than Opus 4.6 at almost everything.
Like 90% of consumer AI use is stuff like "please write a summary of this PDF for me", or "Find a cheaper version of this product online", eventually the returns on model intelligence taper off for these kinds of tasks. However for the kind of frontier tasks they are testing the models on now, like math, science, etc. the marginal returns on increased intelligence are huge.
Anthropic's issues, as I see them, are:
1. They acted as if they were so far ahead capability wise that they could stop listening to their users.
That is basically it. It is an extremely common belief that Opus 5 acts like a condescending wannabe-thought-leader, yet Anthropic's response to this is largely been "You're using it wrong. Try deleting all your config files".
They released a "concise" output format, but pretty much swept it under the rug, despite it being one of the loudest complaints about their models. It also, generally, does not work as described.
Not only did Opus 5 get much more difficult to work with, but it got, from a customer view, significantly slower. The token rate might be the same, but if every interaction takes 50% more tokens, it's 50% slower.
All this time OpenAI has released a slew of new models, lowered the price on them (which, bluntly, 80% of user count don't actually care about since they're on subscriptions), increased subscription capacity, and increased response speed.
It's not about cost. It's about Anthropic being the frontier-lab version of the marathon runner who decides to celebrate to early, and then loses the race.
Opus 5 is by far the best model for everything that matters to me. But it’s just too expensive. Terra 5.6 is bearable for everyday tasks, so now it’s my default.
I have yet to even try Fable or Opus 5. Just looking at the hype they put out prior to the release, then the whole way of releasing these models as well as the pricing just puts me off to get used to it and then needing it. And this while so many good models came out without any hype, botched releases and significantly cheaper. Anthropic really shot themselves in the foot. I went from using sonnet and opus models for 50-60% of my daily token usage to 10-20%
The rate of model improvement has slowed, and may not recover.
It's unclear if Mythos2 or 3 or whatever they're calling their next model will be an improvement for most common enterprise use cases.
LLMs can't solve basic things (writing non-slop documents, understanding context without massive handholding) and for coding other models are quickly becoming 'good enough' without the same cost and nannying.
That's why Anthropic is 'stealing' workflows.
But it turns out it's much harder to push adoption when your users don't really want to use your product.
Code was a unique use case where the code luddites were loud but a minority - most people don't want to update 300 cases of variables across their code base for a name change. Most don't want to write unit tests.
There are a few use cases where that will happen (law is next, maybe quant finance) - but otherwise most companies are throwing money into a pit and getting 0 return.
It's a very interesting race and state of affairs, but Kimi K3 and likely the next DeepSeek models will put the high price token affair to rest.
Unless of course, mythos / next model really does solve some universally applicable problem that people want it it to do.
Perhaps nerfing the cap out of your best model for press attention and hosting valuable features like thought traces isn’t such a great business model?
Are we at the stage yet where a super-strong model that needs oodles of safety protections to stop it outright hacking you is strictly worse for day-to-day tasks than a much cheaper model that's simply not competent enough to be dangerous?
If you are still investing in these companies or plan to in the IPO, the financial ruin you experience is your own doing, you are ignoring every sign that this isn’t going to work out. None of these numbers make sense and point to AI being a commodity with razor thin margins and a race to the bottom. Would you invest heavily in a toilet paper company that took massive amounts of power to produce each version that is 1% softer or stronger every six months?
Every software engineer in my company uses Claude code heavily. However we’ve never enabled Fable and only use Opus, Sonnet and Haiku.
Nobody has complained and seems like for every use case we have Opus is more than powerful enough, especially with Opus 5
> Every software engineer in my company uses Claude code heavily.
I find it funny how OpenAI got caught lacking for a very brief window, but it turned out to be a very critical turning point.
Like a guy that that's at the top of their game the entire year, and the one day they have the flu, the CEO does a surprise performance review.
Yes, the critical point was end of last year, beginning of this year. Especially around the time Opus was released.
i think they were trying to play a different game.
A model is just another thing to plugin to a harness. I don't give it much more thought than that. If developers are still caught up on Claude Code, or Codex that's just not a long term thing. It's best to develop workflows locally and in the cloud with open harnesses. I know this will be the future because that's how it worked on every other system that developers use.
Sure there are Microslop and Oracle db users but most of the world we live in is Postgres and Linux. That's why I think most companies will run llm's like that.
That’s not true. Before AI, I have been using Jetbrains IDEs as far as I can remember. Also have been using MacBooks for work since my first job. You don’t have to generalise everything. If a particular specialised tool is good at its job just use it instead of re-inventing the wheel
Models will be a commodity
> good at its job just use it instead of re-inventing the wheel
Exactly why is everyone reinventing a harness every month. There will be Microslop / Oracle harnesses and there will be 1 or 2 open source ones that win.
I guess soon they will see the full picture
the picture is now clear
I have a $20/mo subscription and have found myself getting limited every few minutes of late. I'm just building a simple website, but after about 15-20 mins of back-and-forth, it tells me I'm tapped out for another 5 hours.
This wouldn't be so bad if it didn't make mistakes periodically, especially when it's about to tap out. I get the sense that if I upgraded to the $200 subscription it would get me a lot more usage, but it would still run into these issues anytime I sat down to work for a few hours.
I'm just using medium effort, so it's not like I'm on high all the time.
I have 200MAX subscription for Anthropic and a RTX 5070TI 16gb GPU. I only use Claude.AI web chat and I only use Opus4.7. I build my prompts on prem with Qwen3.8:latest and copy paste them in claude.ai. I never type anything into claude.ai I am merely a copy pasting monkey. If I need to adjust something I do it from the the on prem prompt. I download code bundles from Claude.AI and push them to my git repo I host. I instruct claude to not do any testing and only let me do testing on my machine so i am not wasting claude resources. The point is to minimize the claude agents from doing anything but specifically doing coding. I can do this for 12 hours a day and reach about 80% of my weekly usage. Once I get to 3 hours left on my weekly I start a fresh chat session with Fable5 and have it do blind code reviews on everything. I only do this because I want to max out my weekly allotment and fable will get me to 20% in 3 hours on a 200MAX sub. Sometimes Fable5 does something interesting but for the most part my on prem static and dynamic code analysis have kept everything buttoned up.
On subscription plan. I only use Fable. It’s been my exclusive coding model since its release. I refuse to use anything else for coding.
For AI agents, we have been mostly using GPT 5.6 Terra.
With open models delivering close to frontier level capability, I really don't see how these labs can only focus on having best model to sustain their business. Although these model's usage will grow, I think and hope that the serving will be distributed among many players.
Omp + 0x alpha is free via openrouter and opencode go
im using fable almost exclusively, i just buy a new 200$ license if i run out of capacity. its so much better, easily worth it given how much work I get done. I built a code, deploy, e2e test loop with my full aws infra, fable just implements linear tasks constantly, 5 at a time, tests the whole thing end to end
if people dont see why they need a model this smart, they probably arent using ai enough
People have realized you don't always need a Fable level model. Majority of my work is sufficient with ChatGPT Luna which has effectively unlimited usage on the $100 ChatGPT plan.
I am not doing awfully complex tasks though. I imagine a lot of other people are in a similar boat, either switching from Claude to ChatGPT or even just min-maxing DeepSeek V4 Flash 0731 or similar.
the code quality is higher though, so the value compounds. i dont need to watch as closely to what its doing, so its more autonomous
Isn't there a risk of getting banned if you do this (multiple accounts to bypass usage limits)?
Idk I never worry about it
Claude will refuse to visit sites with robots.txt, make a graphic spoof of an iOS game, or even fill out an employee survey on my behalf. ChatGPT is always happy to oblige, no questions asked.
I can respect the guardrails - I also can see why OpenAI may not have much control over their models - but I need an AI who will do whatever I ask and not play judge and jury.
Well, as a biologist, Fable is still completely unusable
I can't even ask it how to make toast without zeroing out all its memories of me.
so whats the best and most efficient coding harness and against which model? what are folks doing to keep costs low? I spent $1000 just this weekend on my personal projects for sota Claude but I feel like I can probably get much more juice if I start looking elsewhere.
and resources or tech stack tips from HN?
What a shame, they cooked up all their marketing lies in vain.
Yeah I mean if I run out of tokens every couple of hours and have to pause my work or shell out more money I’ll switch to other tools that don’t have this problem. Though they turned this down a bit it seems, I can work with Fable reasonably now and I enjoy it actually. I think they were just testing out how much they can raise the cost without users leaving when having the best model. I guess not much after all!
I see variations of this post all over X. Fable doesn't get much usage, since Opus 5 is like 99% as good at a lower price point. And perhaps more importantly also faster.
The foundation model companies can’t survive a token price war. If that’s where we’re heading get out your popcorn.
I am getting close to dumping Anthropic. I like their models in general, but boy I have hit their "F you, I ain't gonna help" too many times now on innocent things. Ain't nobody wanna deal with that. I have never gotten that from Antigravity and if I did I woupd tey Codex then go to Openrouter and leave US models behind.
I mean, who says "screw you" to requests to get 35+ year old vintage computers working? Claude, that is who. Its guard rails are so stupid. I hear people trying to do simple mailing list management hit it too.
I am just about done with them.
GLM 5.3 will probably do your vintage computing work without fuss.
Which is why they are rushing to an IPO.
Qwen3.8 is all you need.
People understand the law of diminishing marginal returns.
Anthropic's privacy deviation (ZDR) for Fable is why my organization forbids usage of Fable.
Give me a break, Anthropic is expensive and their CEO is not the nicest guy around, and the games they play with other people's money/their API is not fair. I stopped using Anthropic and OpenAI once they started calling for regulation of open models. I have survived locally since LLama3-70 days and have been surviving fine. If I was to pay for cloud models, it definitely will not be Anthropic, Fable or not. From what I have read, the best AI model will not even comply with requests most of the time because it or/and Anthropic supposedly knows what's better and safe for you.
Do neoclouds win out in this scenario?
Honestly, this makes a lot of sense to me. Devs that are good at their job don't need the absolute best model for every task, and if a cheaper model gets the job done 95% as well, it's a pretty easy choice
"Almost Nobody Is Using Anthropic’s Fable 5" - https://analyticsindiamag.com/ai-features/almost-nobody-is-u...
I'd use it if I could get through a 5 hour session without exhaustion my usage limit.
The market correcting itself.
People here won't believe this, but Anthropic will begin to decline after their IPO when everyone runs to good enough cheaper models to save on token spend.
Not just that but also Opus 5 is real trash. It's free for me (company pays) and I still prefer to use other models.
"If not [bubble]... why bubble shaped?" lol =3
https://www.youtube.com/watch?v=wTiYaWFP59Q
this article reads like a leak from a bank that didn't court the IPO.
Anthropic is the fastest-growing software company in history. They have no issues "attracting users"
I've been saying since the beginning. Models. are. WORTHLESS. If your company depends on having the best model, you have lost.
I've always wondered why everyone flocks to SV's latest darling company. Have we not learned from our history of glorifying these SV darlings that turn hostile?
I think the glib answer is, “Greed blinds all”.
Most of the people pushing this are just hoping that they can cash out before the hype pops and financial gravity crashes the party. Sam Altman recently claiming that the singularity is here is so stupid on its face he should just be treated as what he is, a huckster.
None of this stuff ever made any sense on what it was being sold initially. It was always insulting that the media and business leaders tried to argue that the tech could replace entire call centers or vast swaths of entire industries.
People keep arguing, but it will or it has based on extrapolating certain, reasonable use cases. Klarna has shut up about replacing call centers with bots because Markov chains with memory only can do so much.
Some of us want to get work done and don’t feel the need to either glorify or hypothesize about what might happen.
I’m happy with Claude. If they become (bigger) jerks, I’ll switch to something else. I don’t ha e the energy to praise Anthropic today and I won’t have the energy to demonize them tomorrow. The emotional investment people have for/against these companies feels like celebrating or being offended by the weather.
Are people flocking to them? I see people buy the products but if you ask I think they are just about as hated in the big techs.
A mixture of opportunistic edge-seeking, FUD, FOMO, novelty-seeking, the need to impress shareholders, the tendency of salespeople to believe other salespeople are telling the truth, the ever-present need to stay in front of relentless commodification, and pragmatic curiosity.
TBH I don't think any of that is unique to the IT industry's relationship with the Valley. Other technology-driven industries have a similar worship-ish relationship with a few rarified businesses. But the culture of the IPO exit accelerates all the most short-term motivations to do anything.
Reads like: brilliant Carnegie Mellon University computer science grad struggles to find job where he is not replaced by cheap, inferior Indian labor that still gets the job done, even if it takes marginally longer.
No shit we're all paying 清冲 Flash to do the grunt work. Turns out though, paying 清冲 Flash a few more cents does exactly what Ivy Wasp Pro Mythical does. Crazy how that works.
Absolutely zero people, rounding up generously, are replacing Fable with local or Chinese models.
I am because Claude has figured out I am a biochemist and therefore even asking Fable what the weather is gets me bumped back to Opus, sometimes even Opus 4.8 instead of 5. Kimi? GLM 5.3? DeepSeek? No such problem.
I can literally open a new chat with just "Hello" and it gets bumped.
Memories break Fable 5 for me as well in chatbot. I ask Opus a lot of sec related stuff and now if I even type “hello” in chat it gets insta-downgraded to Opus 5.
Wow, that's totally crazy. Does it happen even if you clear out the "offending" memory entries?
Just cancelled my Claude subscription and used Ox Alpha Free through OpenCode (so that's one).
In my view, from the testing I did since Thursday, it's better than Fable. I had just finished a rather large task that Fable completed, including a /review and an /ultrareview.
Ox Alpha found bugs that Fable and Opus missed, and it continued to build things like a pro.
It does have issues with availability - but it's on a free promo right now. That also means that I don't know how much it would have cost if I had to pay API prices for it, which might not be cheaper than the subsidized small company / consumer usage, but for large companies paying API prices for either offering, the difference will be considerable.
Intriguing. Who is it? What is it? any ideas?
https://openrouter.ai/stealth/ox-alpha
Some people say it is a new version of Gemini Pro - this is based on some tweets from their employees.
Rumored to be mimo
Are you sure, most indicators suggest Z.ai with a potentially Flash or Air variant of GLM 5.3?
I almost feel like there's some sort of paid campaign going on in Hacker News promoting the Chinese/open models. I feel like every day I'm hearing about how the frontier labs are dead but your experience is the same as mine. My company pays for Claude AND Codex but we never really use the open models for anything critical.
Doesn’t matter if your company pays for Claude, Anthropic and OpenAI valuations and expenditures commitment requires them to win the vast majority of the market to make economical sense. And the Chinese competition makes that very unlikely, to say the least. They likely won’t disappear fully but their “free” lunch as the AI darlings is done, on paper. Will be interesting to see how they adapt
Sure, I'm being "paid" by the fact that it costs pennies a day to use Deepseek and if I could afford the hardware then I could run it locally. Meanwhile, Anthropic is obviously a threat to open weight models and actively lobbies the US Government to have them banned or controlled.
So yeah, I and I guess others, are quite active in whatever little way we have available, to up vote new models and share stories.
It's absolutely a campaign. Every time anyone says anything about Claude, immediately and inorganically there's a bunch of people claiming to be biochemists who are constantly shut down by their work, people saying that they run out of tokens instantly even on Premium, people saying they get even better results on their GTX 4050, and people saying that Grok is better, or OpenAI is better. It's probably several different campaigns each run by different tranches of the competition. Reminds me of the good old days in which every criticism of bitcoin was immediately jumped on by nine or ten pretend Venezuelans who asserted that it was the only thing permitting their family to evade government currency controls.
"Everyone that disagrees with me is astroturfing" is not going to lead you to being in touch with reality.
I'm not a biochemist and I have been blocked by fable and opus for "cyber" just for doing things like asking it to ssh into one of my servers, look at a CVE, or do work in assembly. Been rejected to their cyber verification program 3 times already.