The problem I see is the definition of “normal” conversation. To me a normal convo is one I have with humans, not machines.
Graphical User Interfaces are more 'normal' form of interaction than using a computer purely with a command line. It is immaterial whether the system is sentient or not. The relevant factor is whether it is easier to use. If folks don't have to put tons of active thought into how they are talking to the system will be much easier to use. And normal is a good thing. Maximal abnormal is how you make systems harder to use.
Those are the ones I’m concerned about. As the device wearer, you signed up to whatever was in the EULA, they did not. Regardless of location context, whatever someone says to you one on one is implied to be for your ears only unless stated otherwise. If you are in a crowded mall conversing with someone, neither of you expects you might jot be photographed, but what is said to you is still not for everyone at the mall,
What does this have to do with a camera being in the AirPod? The ability to listen in and record conversations is built into all bluetooth headsets that have a microphone attached to them. That is already 'out there'. Same thing with the smartphone with microphones in folks pockets nearby .
Not sure where your are but in USA generally you have to have a reasonable expectation of privacy. A crowded mall , subway car, football stadium is not a context where 'reasonable' secluded privacy. You can hope that folks will be kind and not pay attention, but expectation? No.
and it most certainly is not intended to be used for AI processing or training. Even if Apple were to say (for example, as we don’t know yet) the device ignores X, Y, or Z things, how does it conclude it should ignore them without some processing?
What it doesn't know it won't conclude on. A large extent will be because it doesn't know what some elements are. It is a recognizer which will recognize things that it is trained to resolve.
The part that is more relevant is whether or how deeply it starts to engage in "Phone a friend" to when it does not know.
"...
Other Visual Intelligence Capabilities
Visual Intelligence has a long list of things it can be used for, and these options have been available in earlier versions of iOS.
- Identify plants, animals, insects, landmarks, art, sculptures, books, and more.
- Add items to your calendar from a piece of paper like an event poster or flyer.
- Use Google Image Search to find similar images.
- Search Etsy, Amazon, Anthropologie, and other apps for an item that you capture with the camera.
- Provide details about a business in front of you, like hours of operation.
- Translate, summarize, and read text aloud.
..."
Blinding sending the photo off to 3rd parties user may or may not be aware of would be bad.
If Apple is using this as a perception system ( 'eyes for Siri' ) then there is zero need for the images to be persistent or shared with 3rd parties (who have their own rules). The normal mode for Apple Visual Intelligence is that the photo is put into Photos collection and then go find out what it is. This system really doesn't need to work that way if it part of the perspection sharing between you and Siri at a moment in time. All that needs to be records is what (the context) of what were discussing not the RAW image.
We have already seen that as AI models grow, they often outgrow the capabilities of hardware they may have originally been targeted to, so any notion of “on device” processing doesn’t actually hold water over time. After a couple of software updates,
One reason Apple is 'behind' in the AI mega model race is that is decidedly not the road they took. There are several folks (including Apple) trying to do more with 'less'. The objective is not to build. the largest possible , all knowing, all seeing model. It is to build a model that is useful for a specific purpose. The more specific there is less requirments that is be as large as possible. That is partially because entities like OpenAI, Anthropic can't pull huge valuations unless they are sucking money out of more peoples pockets. One path to that is make folks pay to rent time in their huge datacenters (they may or may not own).
The last 40 years of personal computing has shown that more 'horsepower' has come to smaller devices over time. To read a book title and author off an image doesn't not take a 1 trillion parameter LLM model. OCR was being done years ago on much less capable hardware. Visual inferencing was being over a decade ago on less capable hardware.
if you are asked where images or audio is actually going during a normal conversation, odds are you won’t be able to answer the question with any degree of certainty.
if do zero homework then lacking knowledge comes that process, not from what is knowable.
Apple has talked about what Private Cloud Compute does.
Secure and private AI processing in the cloud poses a formidable new challenge. To support advanced features of Apple Intelligence with larger foundation models, we created Private Cloud Compute (PCC), a groundbreaking cloud intelligence system designed specifically for private AI processing...
security.apple.com
Some data is sent to be inferenced on and then the whole interium dataset is erased.
If Apple optionally punts to the web and/or 3rd parties when Siri doesn't 'know' that would be bad. If they allow Siri to say "I don't know and just stop" then can probably be relatively certain that it didn't go anywhere ( if don't see a snapshot in photos or in the Siri history log. For the latter, Logging the picture as opposed to the metadata would be relatively dumb. Even more so if didn't know what it was. )
Why should anyone have any degree of trust for the device (or its wearer) if they are uncertain of the degree of privacy of any on to one conversation?
Again there is nothing material about the camera aspect here. Airpods with a microphone and the privacy of a verbal conversation are nothing new. If you don't trust the person at all what is to stop them from revealing your conversation to someone with without using a recording device?