ff.aifightfake.ai
← Blog

· Miha Stopar

The reasons for Apple’s Reference Image architecture

Why Apple develops Reference Images in the cloud, and what a capture path without that hop could look like.

My initial reaction when I saw that Apple shipped Reference Image with the iPhone 18 Pro was a relief. Somebody is finally signing the pixels deep in the guts of the device, not somewhere where the signature can be faked. However, a bit surprising fact is that these raw pixels go to the Apple services (Private Cloud Compute or PCC) where they are developed into a photo (demosaicing, tone mapping, and compression).

There goes the privacy then. Apple can see the photos. You might say that if you want to prove the authenticity of the photo, you want to publish it anyway, but that’s not always true: you might want to do some edits before publishing, like hiding the children on the photo.

But it’s true that Apple’s approach has some advantages. For example, in a system where the device signature is the whole proof, if one device gets hacked and the keys get extracted, then these keys can be used freely to make signatures from wherever. The stolen private key signs whatever pixels you like, including AI-generated ones. Apple’s architecture has an additional mechanism to fight this.

In Apple’s case if the Secure Enclave private key was stolen, the attacker would still need the sensor’s signing key. Both are required, because the sensor can be taken out of the phone and put on a rig, so a sensor signature alone does not prove it was still in that iPhone. Apple binds the sensor key and the Secure Enclave key in the factory device manifest, and PCC checks they belong together. The Secure Enclave also signs metadata the sensor cannot see: digital zoom, exposure, and lens parameters. Finally the attacker would need to push the negative to PCC, where a neural net asks whether it looks like real silicon output. PCC runs a neural net with hidden weights that scores whether the negative has the physical characteristics of output from their sensors. Scores accumulate per sensor; a low-scoring sensor can be revoked, and PCC will no longer sign its images. Hidden weights mean an attacker cannot easily train against the detector.

So it might be true that Apple doesn’t trust the Secure Enclave that much, which is a hypothesis by Ian Miers, and they built another mechanism (PCC’s neural networks) to check whether the photo is authentic.

But another reason might be that this actually makes things more secure from another perspective. If you do the development (demosaicing, tone mapping, JPEG) on the phone and only sign after all this process, you put the photo through a pipeline and the attacker has a greater chance to modify the photo during this pipeline. I wrote about that in hardware ramblings and in how difficult can it be to ship a camera that prevents deepfakes: if the hash/signature happens after software has already seen the frames, compromised software can swap in fake pixels. Instead, the sensor signs the photo (the raw pixels) immediately and the phone sends that negative to PCC.

An alternative is to put the image pipeline in the Secure Enclave, then sign the JPEG there, so the OS never sees unsigned pixels and you do not need PCC. But the Secure Enclave is a small coprocessor for keys, biometrics, and signatures. It does not sit on the pixel bus and cannot run Apple’s computational photography stack (demosaic, denoise, tone mapping). That work lives on the ISP (image signal processor: the block that turns the sensor mosaic into viewable frames) and related on-chip imaging hardware.

So I agree with signing before the OS can touch the pixels, but I think there are better options than making Private Cloud Compute the place that develops and vouches for the finished JPEG. What we need is a fingerprint of the viewable frames after the ISP, before software can swap them, plus a signer that will only sign that hardware digest (hardware posts, shippable-camera sketch). You still need a device private key, but it need not be a general-purpose secure element: those chips get broken (firmware bugs, fault injection, side channels). What is better is a small vault welded to the hash path, with no “sign whatever the OS sends” API.