Become a MacRumors Supporter for $50/year with no ads, ability to filter front page stories, and private forums.
Just started using Claude and it sucks hard compared to Gemini and even ChatGPT.

One complex question and I was out of 5 hours worth of "tokens" while Gemini and ChatGPT continued to give more responses for hours.

It's a shame I do like Claude's tone and language but it's frankly useless for me in it's current iteration.
 
  • Like
  • Disagree
Reactions: iMean and Dredd67
Just started using Claude and it sucks hard compared to Gemini and even ChatGPT.

One complex question and I was out of 5 hours worth of "tokens" while Gemini and ChatGPT continued to give more responses for hours.

It's a shame I do like Claude's tone and language but it's frankly useless for me in it's current iteration.
Don't use Fable or Opus 4.8 with High thinking, they absolutely eat tokens. Sonnet 4.6 or Opus 4.7 with Medium thinking are very capable. If you use claude code, you can also turn on the Advisor feature where it will automatically engage a more capable model when it needs too.

If you're on a Claude Pro plan, Opus 4.8 and Fable are probably are not a good fit unless you really know how to manage token use.

On the Google side, has 3.5 Flash gotten better? I know people were hitting limits with it like crazy after it was released at I/O.
 
  • Like
Reactions: tennisproha
It is absolutely wild to me how the EU seems to be still okay with all these not so transparent subscription models of all these AI providers like "may take x credits but we can change it to whatever we want at any given time" isn't very good for the customer.

Also, what's up with all these ridiculous names for each new model. Just call it an update
 
Sorry, dumb questions here. How much money are we talking here:

200-300 million tokens in that time, i.e. not a small amount of use.
How much did those tokens cost in total?
Don't use Fable or Opus 4.8 with High thinking, they absolutely eat tokens. Sonnet 4.6 or Opus 4.7 with Medium thinking are very capable. If you use claude code, you can also turn on the Advisor feature where it will automatically engage a more capable model when it needs too.

If you're on a Claude Pro plan, Opus 4.8 and Fable are probably are not a good fit unless you really know how to manage token use.
How much money does it take to do -- what?

Is this being used for code conversion, for example? Suppose you wanted to convert 10,000 lines of C++ to Rust? How many tokens is that and what would it cost? As I said, dumb questions, but, I don't really understand what you folks are talking about. 😱
 
I'm more worried about the people who have access to the full unfettered model.

We're already at the point where these things are de facto "military grade."

Cybersecurity, chemistry, biology. Interesting that cybersecurity is now mixed in with the world class deadly things to know too well.

I hang around in cybersecurity circles. The reality is this is entirely overblown. While there are some gains the main problems are actually supply chain attacks, which really are just poor methodological and architectural design issues. Oh and the fact this has caused a gold rush on defects. Capability-wise the gains are minimal. There are skilled folk who run rings around these models. But you wouldn’t know it until you’ve worked with them.

I’ll add that my daughter is a biochemist and thinks the proposed use cases there are mostly fantasy too.

There’s a huge amount of hype.
 
Last edited:
Sorry, dumb questions here. How much money are we talking here:


How much did those tokens cost in total?

How much money does it take to do -- what?

Is this being used for code conversion, for example? Suppose you wanted to convert 10,000 lines of C++ to Rust? How many tokens is that and what would it cost? As I said, dumb questions, but, I don't really understand what you folks are talking about. 😱

This one always gets me.

It’ll be cheaper than humans but the only reason it gets done is because this tool is available. And when it is done, if you like it or not, porting to Rust doesn’t necessarily make it better or more secure. The coreutils rewrite in Rust was a fine example of this. More CVEs than the originals C version. Idiocy everywhere.
 
How much did those tokens cost in total?
Some of that wasn't well tracked, I've had a few long repeatedly compacted many day sessions with usage in the ~8k range so probably about ~$60k total I'd guess. I tend to start fresh and keep things smaller these days.

The few running instances I have right now, all started today, are about $300-400 combined total right now to give you an idea.

This stuff is NOT cheap using APIs. But the work I'm doing is fairly complicated, I'm not vibe coding a webpage or basic app I'm doing research (and this is also why I keep hitting these freaking classifiers, so annoying).

To give you an idea one series of things I was chasing for 2 months turned out over 600 python scripts to test various permutations of experiments that I needed to prove-out theories. This sounds like inefficiency but it really wasn't, my velocity otherwise would have taken years and I was able to disprove something I'd otherwise have chased for a long time, but I got some useful things out too.

I am absolutely dying for an M5 Ultra, I'm constantly out of RAM or CPU / GPU constrained on my M4 Max.


I also do NOT use agents except to search for things either in my history (e.g. pull up this facet of my research work across many documents and synthesize for a new session) or broader web search tasks. The sub-agent workflows are not worth the money and produce sub-par outputs. Ultraplan is kind of a joke, but max effort sure isn't.

I need to experiment more with that later this summer, I'm going to try some workflows with codex but I prefer Claude by a wide margin. I use codex in zed sometimes though, and probably would use it in Xcode also if I developed more for Apple, which might happen given you can now write Kernels with MLX as of this WWDC but I need to look into this more and it seems like kind of a distraction I don't need at the moment.
 
Last edited:
Don't use Fable or Opus 4.8 with High thinking, they absolutely eat tokens. Sonnet 4.6 or Opus 4.7 with Medium thinking are very capable. If you use claude code, you can also turn on the Advisor feature where it will automatically engage a more capable model when it needs too.

If you're on a Claude Pro plan, Opus 4.8 and Fable are probably are not a good fit unless you really know how to manage token use.

On the Google side, has 3.5 Flash gotten better? I know people were hitting limits with it like crazy after it was released at I/O.
I switched from Claude to Gemini for this exact same reason. Why pay 20 bucks with only being able to get a glimpse of half the features before getting slapped with an abrupt "hands off dude, you're too greedy, you've already asked ONE question today".

My experience is that I'm using gemini 3.1 pro most of the time without hitting any limits. I ask complex questions, involving documents, code, cross analysis, sometimes canvas (artifacts). Only once did I get a downgraded answer by 3.5 flash, since 3.1 was in high demand. Asked again, 3.1 pro answered. So no biggie here.

Never hit limits for 3.5 flash either.

I tried ChatGPT, but also hit the limit a few times.

This video from futurepedia is a good analysis for most people :

 
Claude Fable 5, a Mythos-class model that it says is "safe for general use"...

In the rapidly growing AI space where anything goes, funding flows like water and the narrative is tightly controlled, it's buyer beware. Those in the public sector who get drawn in believing the hype while failing to read the terms and conditions are fueling the madness. While certainly there is some usefulness, the opposite could also be argued. The word that comes to mind is nefarious.
I sent Anthropic question whether they tested Fable using Mythos to be sure its security protocols wont be hacked by some low level hacker using some other AI 😆
 
  • Haha
Reactions: zenmacx
I hate the obscure pricing models these companies offer. It is jsut like mobile internet access in the early days. About time to normalize and drop in price if you ask me. And get rid of the ridiculous token counting, that's internal tech stuff, not exactly useable for the average user.
 
  • Like
Reactions: mistymac
This article is written like I know what any of this means (spoiler: I don’t). Mythos? Fable? Opus? Input/output tokens? Also, I don’t see how any of this is related to Apple news or rumors.

I agree, pardon my ignorance, but if you are going to write an article please give it some context for the lay people out there. Took me several searches to find out that you were actually talking about AI 🙄🙄 when I realised what it was about I lost interest....
 
Yes and it will be passed up in a few weeks by Gemini and Open Ai and in a year by Grok (the slow kid). We are in the phase of rapid model improvement. This is how it works.
 
Apple put out a video yesterday on how to hook up a local model to Xcode. So I’d say, yes, they are promoting vibe coding.
Using an LLM to generate code is not "vibe coding". The term is completely misunderstood. Vibe coding is someone who cannot understand/write/audit code, asking an LLM to generate code and using that code verbatim, without being able to understand it.

An experienced software engineer asking an LLM to generate boilerplate code, fix a bug, write a feature, etc. and then — crucially — auditing that code to ensure it fits with the larger codebase, is not vibe coding. It's efficiency.
 


Anthropic today announced the launch of Claude Fable 5, a Mythos-class model that it says is safe for general use.

claude-fable-5.jpg

According to Anthropic, Fable 5's capabilities exceed those of any model it has made generally available, and Fable has demonstrated "exceptional performance" for software engineering, knowledge work, vision, scientific research, and more. It outperforms Opus models on longer, more complex tasks. Fable 5 can work autonomously for longer than any prior Claude model.

Fable 5 is being released with conservative safeguards to prevent it from being misused in areas like cybersecurity. Questions about some topics will instead be answered by Opus 4.8, with safeguards expected to trigger in less than five percent of sessions on average. Most queries related to cybersecurity, chemistry, and biology will get responses from Opus 4.8 instead of Fable 5.

Anthropic is also releasing Claude Mythos 5 for a small group of cyberdefenders and infrastructure providers. It uses the same underlying model as Fable 5, but with some of the safeguards lifted. Mythos 5 is being deployed through Project Glasswing as an upgrade to the Claude Mythos Preview. Anthropic says Mythos 5 has the strongest cybersecurity capabilities of any model in the world, with access set to expand through a broader trusted access program.

Fable 5 and Mythos 5 are available at $10 per million input tokens and $50 per million output tokens, which is less than half the price of the Claude Mythos Preview. Mythos 5 is available to those who have access to the Mythos Preview, and that includes Apple. Apple is one of Anthropic's Project Glasswing partners.

Claude Fable 5 is included in Pro, Max, Team, and seat-based Enterprise plans from today until June 22. On June 23, the model will be removed from those plans and using it will require usage credits. When Fable 5 capacity is sufficient, Anthropic plans to re-add it to subscription plans.

Article Link: Anthropic Launches Claude Fable 5, Its First Public Mythos-Class Model
I should warn you all. I tested this - and it's context is just 240k... so Opus still out performs it in understanding context and tackling large research.

Sorry, Fable 5 didn't work very well for my large code set.
 
Don't use Fable or Opus 4.8 with High thinking, they absolutely eat tokens. Sonnet 4.6 or Opus 4.7 with Medium thinking are very capable. If you use claude code, you can also turn on the Advisor feature where it will automatically engage a more capable model when it needs too.

If you're on a Claude Pro plan, Opus 4.8 and Fable are probably are not a good fit unless you really know how to manage token use.

On the Google side, has 3.5 Flash gotten better? I know people were hitting limits with it like crazy after it was released at I/O.
I use Gemini Pro but I find AI Mode works better with some deep analytical stuff.
 
Register on MacRumors! This sidebar will go away, and you'll see fewer ads.