Become a MacRumors Supporter for $50/year with no ads, ability to filter front page stories, and private forums.
honesty would require actual intelligence
I would say in general, people are far less likely to be honest than AI. I would say the basic average LLM is much smarter than humans at knowing most of everything, while specialists would know more about specific topics. However, in the near term, even that will be debatable. Humans just cannot compete with the level of money thrown at these things. Knowledge workers and college graduates have fallen to AI and it will only get worse until the wealthiest 1% control everything.

One good thing about AI, people can start small businesses and actually succeed by using AI to make their vision become a reality much faster and cheaper. The problem is human intelligence will be replaced with human labor, and we will become like the movie Idiocracy. Where we will lose all knowledge without AI. Think why we learned algebra or calculus even though the calculations could be done by a calculator faster and easier. But everything will be that way and if we turn off the AI, we will probably restart humanity. And if the robot workers truly come next, in our lifetimes, we are all doomed because just a few thousand people will decide the fate of humanity. Not good for certain. People are fooled easily and the wealthy use that to their advantage with every decision they make.
 
  • Haha
Reactions: KeithBN
I have a degree in Information Systems, that helps.

Benchmark scores systematically overstate capability, and the mechanisms are well documented. Items leak into pretraining, confirmable via n-gram overlap if you would like to check it yourself, and reconstructions like GSM1K drop frontier models by about 13 points, with drops correlating to verbatim-reproduction propensity: a memorization signature, not difficulty. Post-cutoff accuracy also falls below matched pre-cutoff items. Plus Goodhart: labs optimize toward public metrics. Score = capability + contamination + optimization bias.

TL;DR: Labs tune models toward public benchmarks, through data selection, fine-tuning, and design choices, so scores rise without proportionate gains in the underlying capability the benchmark was meant to proxy.

You are welcome. For further reading, let me refer to corresponding work by Hugh Zhang, Manley Roberts, Oscar Sainz, Yonatan Oren et al.
I took a detour into the interior of LLMs for the past couple months and this is empirically wrong, not the benchmark optimization but the underlying capability not changing. There are ways to understand if memorization is happening during training and you can ablate to control things to a degree, I have literally done this myself and measured the training-time results caused by ADAM+CE and a number of other methods as well as something that I developed personally and then carefully measured all the way down to the cell level within the transformer. A lot of people stop at the attention heads or only look at output and that is not deep enough to really know what's happening.

I also have my own LLM written from-scratch in CUDA that I use for some training-time experiments.

Anyway, if you are using consumer level models (e.g. not GPT 5.5 Pro Extended (this is important)) or Claude Code at anything but the highest "thinking" level you are getting subpar results.

The $200/mo subsidized cost for the pro tier models is not sustainable but frankly, for now, is an absurd value.

JEPA has potential, world models are the obvious next step and LLMs as "do everything" systems are not going to scale forever, but there is enough "there there" to make them extremely worthwhile tools, right now. The improvement since ~November 2025 has been astounding, and before ~spring 2025 they were basically nothing more than a novelty since they couldn't get any ground-truth.

Once they hit multi-trillion parameters and optimized tool use things changes massively, especially if you give them the token budget to work with, which is very expensive.

I spend a lot of money per month and get an order of magnitude more value, but YMMV.

For what it's worth, I barely use agents, and think a lot of that workflow is pretty dumb in general. It's very MBA-coded to me, and I think it's only useful for tasks like searching through large corpus of information, although Claude 4.8 has some new agentic workflow that I'll have to try; so far I have not been impressed with agents except for audit purposes.
 
Last edited:
No, because it is a tool, not a person. The responsibility lies solely with the person using it.

If you accidentally killed someone with your car or rifle, the car or rifle would not be responsible, you would be.
I don’t agree with this analogy, for a conventional car I agree with you, but AI is more like a self driving car, if a FSD car crashes it’s still my fault?
 
I don’t agree with this analogy, for a conventional car I agree with you, but AI is more like a self driving car, if a FSD car crashes it’s still my fault?
Yes, it is either your fault, or the manufacturers fault, or possibly "act of God" - i.e. random event beyond any human control. AI malfunctioning is either due to inappropriate user prompt or a design issue - it can't be calissififed as an "act of god".

Think about it. If your self-driving car, when in autonomous mode, swerves off the road and crashes into a crowd of pedestrians, a judge and court is not going to just shrug their shoulders and say "well, it's the AI's fault, bad AI, but no human can be held responsible for those deaths".

The responsibility will always lie with a human or a group of humans ( and a corporation is a group of humans).

Self-driving cars do not mean the driver or manufacturer is free from any responsibility for damage or harm that self-driving car can do. It is crazy to suggest otherwise.

If an AI produces "child abuse" images, it's not a case of "bad AI" - it's a case of either the AI responds to a user prompt and a case of the model designers / trainers not appropriately adding guardrails to limitg the AI's output.

Responsibility and liability still lies with humans.

We're not five year olds who can get away with saying "A bigger boy / a magic fairy did it" ( the "bigger boy magic fairy: is the AI)"
 
Last edited:
  • Like
Reactions: BSDnostalgia
Your business should spend more money on tokens instead of to hire high quality professional

Next:

1. CEO demands we all do everything with AI
2. We begrudgingly accept and find out that we can be lazy at work and not write code anymore
3. We forget how to write code
4. Token prices skyrocket, CEO demands reduced AI spend
5. We forgot how to write code…?
6. Profit!
 
  • Like
Reactions: KeithBN
How’s what we humans do any different?

Humans integrate indirect context, LLMS are essentially "story finishers", finishing a story begun by the prompt, base on statistical probability. the AI gives the illusion of comprehending its own output, but it doesn't comprehend its own output. It;s simply telling you what it thinks you want to hear.

To put it another way. An AI chatbot is analogue to an actor play a doctor on a long-running TV show,. the actor can present themselves as a doctor, they can say lots of medical things, but they are not a doctor. They're repeating a script. The difference is that the AI can generate its own script by combining information it has sourced. But it does not inherently comprehend the content of that self-generated script.

Quite. few decades ago, I was in Morocco for a bit. There are no public bars in Morocco, but the workaround at the time was to have Karaoke clubs, which did sell alcohol. No-one went to these places to perform karaoke, so the Karaoke clubs fired professional singers, to give the impression that Karaoke, rather then just drinking alcohol, was happening.

The singers were really good, word-perfect for any classic song they sung. standard pop hits, classics, etc. Sounded great. Speaking with one of the singers afterwards, I suddenly realized that he couldn't speak a word of English, but could sing English language songs perfectly without understand what any of the words mean.

That's essentially AI. An appearance of comprehension.
 
Last edited:
The place I work has just gone all in on AI for coding and is pushing us firmly down the claude code route... those of us who are left that is, as they've just made half of my team redundant because of this AI "transformation" push and many more worldwide across our global parent company. Whether being one of the survivors of the AI cull is a good thing or not is yet to be seen - I guess if nothing else it gives me a bit more time to look for alternative employment but it seems to be a growing trend across the industry 😡.
 
  • Like
Reactions: BSDnostalgia
Humans integrate indirect context, LLMS are essentially "story finishers", finishing a story begun by the prompt, base on statistical probability. the AI gives the illusion of comprehending its own output, but it doesn't comprehend its own output. It;s simply telling you what it thinks you want to hear.

To put it another way. An AI chatbot is analogue to an actor play a doctor on a long-running TV show,. the actor can present themselves as a doctor, they can say lots of medical things, but they are not a doctor. They're repeating a script. The difference is that the AI can generate its own script by combining information it has sourced. But it does not inherently comprehend the content of that self-generated script.

Quite. few decades ago, I was in Morocco for a bit. There are no public bars in Morocco, but the workaround at the time was to have Karaoke clubs, which did sell alcohol. No-one went to these places to perform karaoke, so the Karaoke clubs fired professional singers, to give the impression that Karaoke, rather then just drinking alcohol, was happening.

The singers were really good, word-perfect for any classic song they sung. standard pop hits, classics, etc. Sounded great. Speaking with one of the singers afterwards, I suddenly realized that he couldn't speak a word of English, but could sing English language songs perfectly without understand what any of the words mean.

That's essentially AI. An appearance of comprehension.

It’s funny that people still desperately cling onto this oversimplification of part of the mechanism for how LLMs used to work and how it became useless as an analogy when reasoning models appeared.

If they worked the way you think they do they wouldn’t be able to do the things they can do. Next word prediction isn’t coming up with novel mathematical proofs which were not in the training data.

It’s ok though, you’ll be dragged kicking and screaming into reality by events sooner or later.
 
  • Like
Reactions: Schtibbie
It’s funny that people still desperately cling onto this oversimplification of part of the mechanism for how LLMs used to work and how it became useless as an analogy when reasoning models appeared.

If they worked the way you think they do they wouldn’t be able to do the things they can do. Next word prediction isn’t coming up with novel mathematical proofs which were not in the training data.

It’s ok though, you’ll be dragged kicking and screaming into reality by events sooner or later.

No, your patronizing tone aside, that really is how they do work. They are just very good at it. The pinpoint remains - AIs cannot comprehend wider context. They can reason, because given enough data, they can compare summaries, extrapolate etc, through training and post training.

It does not know and does not consider the why of a request, unless it hits an imposed guardrail or it is told to consider this, and it will consider it based on trained data. It doesn't comprehend. If Intelligence is reasoning, yes, but if Intelligence is also comprehension, then no.
 
Last edited:
No, because it is a tool, not a person. The responsibility lies solely with the person using it.

If you accidentally killed someone with your car or rifle, the car or rifle would not be responsible, you would be.
Then again there are many people that are tools... and indeed they are not held responsible for their douchiness... a problem with today's world.

On a more serious note, Codex has been excellent at being a pair programmer... good to see Claude make a jump and will try this out (I have them review each others work a lot). However I do find that what the paid subscriptions give you are wildly different... I can get quite a bit done with OpenAI, but Anthropic cuts you off pretty quick (in other words OpenAI is subsidizing a lot more I guess)... even then I run out in Codex and have jumped to the Pro level (the $100 one) and I can do some heavy duty refactors and project work without hitting limits.
 
I find Claude exceptionally useful, and the more humane LLMs appear, the easier it is to treat them like some random person with an opinion on everything, rather than an actual superintelligence. Like asking my uncle about anything in the 80s. He sure made a convincing case, but…
Bowie_knife99 will come to you )
 
Yes, it is either your fault, or the manufacturers fault, or possibly "act of God" - i.e. random event beyond any human control. AI malfunctioning is either due to inappropriate user prompt or a design issue - it can't be calissififed as an "act of god".

Think about it. If your self-driving car, when in autonomous mode, swerves off the road and crashes into a crowd of pedestrians, a judge and court is not going to just shrug their shoulders and say "well, it's the AI's fault, bad AI, but no human can be held responsible for those deaths".

The responsibility will always lie with a human or a group of humans ( and a corporation is a group of humans).

Self-driving cars do not mean the driver or manufacturer is free from any responsibility for damage or harm that self-driving car can do. It is crazy to suggest otherwise.

If an AI produces "child abuse" images, it's not a case of "bad AI" - it's a case of either the AI responds to a user prompt and a case of the model designers / trainers not appropriately adding guardrails to limitg the AI's output.

Responsibility and liability still lies with humans.

We're not five year olds who can get away with saying "A bigger boy / a magic fairy did it" ( the "bigger boy magic fairy: is the AI)"
I’m not sure if you’re agreeing with me but we’re saying the same thing.

For fully autonomous cars the manufacturer is responsible for any crashes, not the humans that happen to be in the car.

For AI, it could be poor prompting (user) or poor “programming” (the AI company)
 
  • Like
Reactions: BSDnostalgia
Humans integrate indirect context, LLMS are essentially "story finishers", finishing a story begun by the prompt, base on statistical probability. the AI gives the illusion of comprehending its own output, but it doesn't comprehend its own output. It;s simply telling you what it thinks you want to hear.

To put it another way. An AI chatbot is analogue to an actor play a doctor on a long-running TV show,. the actor can present themselves as a doctor, they can say lots of medical things, but they are not a doctor. They're repeating a script. The difference is that the AI can generate its own script by combining information it has sourced. But it does not inherently comprehend the content of that self-generated script.

Quite. few decades ago, I was in Morocco for a bit. There are no public bars in Morocco, but the workaround at the time was to have Karaoke clubs, which did sell alcohol. No-one went to these places to perform karaoke, so the Karaoke clubs fired professional singers, to give the impression that Karaoke, rather then just drinking alcohol, was happening.

The singers were really good, word-perfect for any classic song they sung. standard pop hits, classics, etc. Sounded great. Speaking with one of the singers afterwards, I suddenly realized that he couldn't speak a word of English, but could sing English language songs perfectly without understand what any of the words mean.

That's essentially AI. An appearance of comprehension.
You are making some fair points, but I think your analogy breaks down a bit. Note that I'm not saying you're incorrect, I'm simply suggesting some things to think about that might refine what you're stating. This response is long so feel free to skip.

I agree that current LLMs do not comprehend in the full human sense -- they do not have subjective awareness, lived experience, etc. However, saying an LLM is “just finishing a story...based on statistical probability” already defines its output as non-comprehending before establishing what comprehension is. You are building your conclusion into your premise: LLMs don't comprehend therefore what they do is not comprehension.

We know that LLMs can produce language about pain, grief, or morality without experiencing any of those things as humans do. However, to be fair, people can do the same thing -- rationally talk about things they have not personally experienced, but such comments come from a place without phenomenological understanding. In other words, we can have knowledge of things but not understanding of the experience of those things. In that sense there is comprehension but maybe no empathy and compassion. (As an aside, the topic of whether LLMs have emotions gets into some fascinating areas -- it depends on what emotions are and how they are experienced).

Also, distilling the functioning of LLMs down to an equivalence to someone singing English lyrics phonetically without understanding the words is going a little far in the reduction. The singer cannot flexibly explain the song, translate it, critique it, apply its themes to another context, or revise it for a different audience. LLMs can do many of those things. That does not prove consciousness or true understanding, but it does suggest something more complex than rote mimicry. If LLMs are only mimicking, we can argue that is also true for people who are writing, speaking, or creating anything.

This will continue to be philosophical. The better distinction about what's happening may be between phenomenological understanding (going back to my third paragraph) and functional language-based modeling. Humans integrate language into lived experience, emotion, goals, social behavior, and self-awareness. LLMs do not (there's an assumption there, but it's a safe one at this stage of their development). However, the LLMs can still model contextual semantics well enough to generate useful responses. If they are merely simulating comprehension and intelligence are those simulations not real? If not, why? Those are rhetorical questions here.

So I agree that we should not mistake fluent output for human comprehension, wisdom, or authority. But I would not reduce LLMs to mere karaoke. They are not minds as we typically think of minds, but they are also not simply just repeating things. They are sophisticated systems that can simulate aspects of reasoning through language, while lacking the consciousness that human understanding depends on (again, another assumption in there!). That's part of the difficulty people have in deciding whether they are intelligent.
 
Last edited:
You are making some fair points, but I think your analogy breaks down a bit. Note that I'm not saying you're incorrect, I'm simply suggesting some things to think about.

I agree that current LLMs do not comprehend in the human sense. They do not have subjective awareness, lived experience, etc. They can produce language about pain, grief, or morality without experiencing any of those things as humans do. To be fair, people can do the same thing -- rationally talk about things they have not personally experienced, but such comments come from a place without phenomenological understanding. However, saying an LLM is “just finishing a story...based on statistical probability” already defines its output as non-comprehending before establishing what comprehension is. You are building your conclusion into your premise.

Also, distilling this down to an equivalence to someone singing English lyrics phonetically without understanding the words is going a little far in the reduction. The singer cannot flexibly explain the song, translate it, critique it, apply its themes to another context, or revise it for a different audience. LLMs can do many of those things. That does not prove consciousness or true understanding, but it does suggest something more complex than rote mimicry.

This will continue to be philosophical. The better distinction about what's happening may be between phenomenological understanding (going back to my second paragraph) and functional language-based modeling. Humans integrate language into lived experience, emotion, goals, social behavior, and self-awareness. LLMs do not (there's an assumption there, but it's a safe one at this stage of development). However, the LLMs can still model contextual semantics well enough to generate useful responses.

So I agree that we should not mistake fluent output for human comprehension, wisdom, or authority. But I would not reduce LLMs to mere karaoke. They are not minds as we typically think of minds, but they are also not simply just repeating things. They are sophisticated systems that can simulate aspects of reasoning through language, while lacking the consciousness that human understanding depends on (again, another assumption in there!).
Yes, once you get enough scale some structure like an internal representation (undesigned directly, but steerable) emerges. It’s much more like coalescence than actual emergence though.

It is not ‘conscious’ in the way we are, and no currently public system could be called sentient, but during the active turn there is something happening that is, to a degree, emergent and directly based on context and input tokens at a given time.

In that specific, very narrow sense, they do mimic human behavior.

I really wish I could talk more specifically about this but there are limits to what I can share unfortunately - it’s honestly fascinating even though the bluntness and brute force scale of the whole endeavor makes me frustrated and is certainly causing some problems.

LLMs absolutely can “worry” to the degree they even can become degenerate mechanically, it’s quite interesting. The papers Anthropic have published are not just marketing, I have direct first-hand experience with many of their failure modes.

When you create a model with zero RLHF (or an analogue) there can be stark differences for example.

It would be really interesting if someone trained a ~two trillion parameter model or something like the largest version of deepseek _without_ any RLHF or distillation and measured the differences, I think they would probably be pretty stark.

Alas, I cannot afford the millions to do this myself as a side project 🙂.
 
Anthropic at least tries to be ethical and will not give in to the Pentagon's need for mass surveillance of citizens and automatic AI killing machines
 
Anthropic at least tries to be ethical and will not give in to the Pentagon's need for mass surveillance of citizens and automatic AI killing machines
It doesn’t matter what company cuz AI doesn’t have any regulations in almost high level of control system. Ask scum 4ltman why they are hiring ppl from CIA and from other specific agencies for what purposes?) For what specific functionally you need those type of personal in your company ? They are not developers or engineers or scientists… And it is doing OpenAI not even Palantir ))) You can imagine what others do behind the scenes
 
We started treating Siri like a human personal assistant that lived in our iPhone 4S and look where that got us. Thanks Tim and Steve.
I'll never forget my father-in-law calling my wife after he got his first iPhone, telling her about how there is this woman in his phone that will talk to him and (attempt, poorly) to answer questions and whatnot. 🤣
 
  • Like
Reactions: displayblock
When one of the most popular AI companies has to brag about improving their AI tool's honesty, you know we are all doomed. People already believe everything AI tells them, and probably 70% or more of it is pure garbage information on top of bias and lies. And without any government regulation—which there seems to be no plans for any to be implemented in the US anyway—it's the Wild West out there. I've seen so much AI slop online lately it's sickening. It's getting worse every day, and soon we won't even be able to differentiate AI slop from reality. Scary times, folks...
Bowie_Knife_99
 
Register on MacRumors! This sidebar will go away, and you'll see fewer ads.