# ANE Compile Cost and Compiled-Bundle Caching for Large Core ML Models

> **Collective Library edition.** This is the complete technical report. Private filesystem paths, internal run identifiers, campaign-control notes, and repository navigation were removed. Technical claims, code, measurements, evidence labels, citations, corrections, and falsification criteria are preserved.

Why a large Core ML encoder can take twenty minutes to load the first time,
where the compiled artifact goes afterwards, and how to prove the Neural Engine
was actually used.

**Provenance.** Ingested 2026-08-07 from a Gemini Deep Research Max run:

```text
  apple-neural-engine-compilation-cost-and-compiled-bundle-caching-for-large-core-/
  2026-08-08T03-52-03-618Z/raw-report.md

SHA-256 f3b30252b62ac1f06f139735234e267fc3ae646d3394a0e3725689bf76aa6695
```

Compute-unit and `MLComputePlan` semantics re-checked against Apple's Core ML
documentation via Context7. Sections marked **Implementation evidence** are firsthand
measurements from the reference implementation and take precedence over the draft where they
conflict.

## Executive summary

- **A cold ANE compile of a ~1 GB transformer encoder takes 10-50 minutes on
  M1**, and practitioners report this as normal rather than pathological. Cost
  is driven by graph segmentation and fallback insertion, not parameter count.
- **The compiled bundle is cached per calling executable**, under
  `~/Library/Caches/<executable-name>/com.apple.e5rt.e5bundlecache/`. The cache
  key is a digest over the MIL program, bound to the OS build and the caller's
  cache domain.
- **`MLComputeUnits.all` is not a directive to use the Neural Engine.** Apple
  documents it as letting the OS choose; it can load fast by scheduling around
  the ANE entirely. Only `MLComputePlan`, an Instruments Core ML trace, or ANE
  power in `powermetrics` proves placement.
- **`ANECompilerService` is a single shared root-owned queue.** Spotlight
  indexing and Photos media analysis saturate it, and any ANE model load then
  waits behind them.
- **The ANE has no general matrix-multiply unit.** It executes matmuls as 1x1
  convolutions, which is why `einsum`-heavy attention graphs compile slowly and
  why Apple's `ane-transformers` rewrites attention into 4D `Conv2d`.

---

## Compiled-bundle caching

### Location and cache key

Compiled ANE artifacts ("E5 binaries", after Apple's Execution Engine 5 runtime)
are written to `~/Library/Caches/<identifier>/com.apple.e5rt.e5bundlecache/`.
Below that the layout is `<os-build>/<hash>/`, where the hash is a digest over
the MIL program. The cache key therefore includes the model graph and weights,
the exact OS build, and the caller's cache domain.

System daemons get their own domains — media analysis caches under its own
container, Spotlight under `com.apple.Spotlight`.

Because the key includes the OS build and because ANE topology differs across
Apple Silicon generations, **a bundle cannot be pre-built by a developer and
shipped**. The compile happens on the end user's machine.

> **Implementation evidence — the identifier is the executable name, and the shared
> directory is a red herring.** Measured on an M1 Mac mini running macOS 15.7.7:
>
> ```text
> ~/Library/Caches/cu/com.apple.e5rt.e5bundlecache               4863 MB
> ~/Library/Caches/encoder-bench/com.apple.e5rt.e5bundlecache    3646 MB
> ~/Library/Caches/kokoro-worker/com.apple.e5rt.e5bundlecache    9014 MB
> ~/Library/Caches/whisper-cli/com.apple.e5rt.e5bundlecache         1 MB
> ~/Library/Caches/com.apple.e5rt.e5bundlecache                    29 MB
> ```
>
> `cu` and `encoder-bench` are bare command-line binaries with no bundle
> identifier, and each got a cache domain named after the executable. The
> unqualified `com.apple.e5rt.e5bundlecache` at the top of `Caches` is barely
> used. **Rename or rebuild a binary and it starts from a cold cache**, because
> the domain follows the executable name.

### Size ceilings and cache failure

The draft attributes multi-minute recompiles on every load to an undocumented
~2 GB ceiling, citing practitioner reports of
`BNNS Graph Compile: failed to preallocate file with error: No space left on
device` when a compiled bundle grows past roughly 2 GB, with jetsam killing the
process or the write being abandoned. Reports of that error are real and worth
knowing.

> **Implementation evidence — the "silently fails to persist" story did not hold here,
> and believing it cost hours.** A 1.2 GB Whisper `large-v3-turbo` encoder
> compiled for `cpuAndNeuralEngine` in **985-1300 s** cold, and the identical
> load with the bundle cached took **0.64 s**. Caching worked. The apparent
> "nothing persists" symptom came entirely from watching the wrong directory:
> the unqualified `com.apple.e5rt.e5bundlecache` never grew because the
> compiler was writing to `~/Library/Caches/cu/…` and
> `~/Library/Caches/encoder-bench/…` instead. Two published conclusions were
> retracted over this. Check the per-executable domain before concluding a
> persist failure.

### Eviction

Cached bundles are volatile. OS updates, low disk space, and long disuse purge
them, so a load that was fast last month can be slow today. Applications that
care about first-inference latency prewarm rather than assume residency.

> **Implementation evidence — treat a cache directory as load-bearing infrastructure.** A
> speech worker on the same host held 9 GB of ANE bundles. Clearing
> `~/Library/Caches` on that machine, or shipping a renamed worker binary, would
> discard all of it and force a full cold recompile on the next model load.

---

## What makes ANE compilation slow

### Graph segmentation and fallback insertion

The compiler tries to place every operation on the Neural Engine. When it hits
one the ANE cannot execute — because of precision, dynamic control flow, or
tensors exceeding ANE SRAM — it must cut the graph and insert a transfer to CPU
or GPU. On large graphs this produces hundreds of "islands"; one reported case
inserted 526 ANE-to-CPU/GPU context transfers and 262 single-op fallback
segments. Compilation passes then do work proportional to layers times
transfers, which is where tens of minutes come from.

Compile cost tracks **graph structure**, not parameter count. A model that maps
cleanly compiles far faster than a smaller one that does not.

### `einsum` versus `Conv2d`

The ANE has no general matrix-multiply unit. It executes matmuls through its
convolution datapath as 1x1 convolutions. `einsum` is valid Core ML MIL, but an
attention graph built from many `einsum` ops has to be decomposed into
convolution tiles or handed to CPU fallbacks for layout reasons.

This is why Apple's `ane-transformers` reference rewrites attention into 4D
`Conv2d` form and removes `einsum` entirely. The ANE also wants strict stride
alignment — the last axis a multiple of 64 bytes — and unaligned tensors force
the compiler to insert layout-conversion layers.

> **Implementation evidence — a stock Whisper encoder export is `einsum`-heavy.** The
> compiled `large-v3-turbo` encoder used here reports
> `Einsum 1280, Softmax 640, Conv 194, LayerNorm 65` in its MIL op histogram,
> and the same histogram appears in the artifact produced by whisper.cpp's own
> converter. **Both took 985-1300 s to compile cold**, so the long compile is a
> property of this model shape rather than of any one exporter. Whether the
> `einsum` count is *causal* here was not tested; that would need
> `MLComputePlan` op-level placement, which is the check to run before acting on
> this hypothesis.

---

## Compute units and silent fallback

Apple documents the cases plainly: `all` lets the OS choose the best option
across CPU, GPU, and Neural Engine and is the default; `cpuOnly`, `cpuAndGPU`,
and `cpuAndNeuralEngine` restrict it (verified via Context7 against Apple's
`MLComputeUnits` documentation).

The practical consequence the draft draws out: **`all` is permission, not a
requirement.** If ANE compilation is expensive or the graph maps poorly, the
planner can schedule the model onto GPU/CPU and return quickly, while
`cpuAndNeuralEngine` forces the expensive ANE path. A fast load under `all` and
a twenty-minute load under `cpuAndNeuralEngine` for the same artifact are
consistent with this.

> **Implementation evidence — the fast-load-means-`all` explanation does not fit one
> important local case, so treat it as a hypothesis.** The draft proposes that a
> previously recorded 465.9 ms load in the reference implementation must have been `all`
> quietly choosing the GPU. But the build that produced that measurement was
> patched to request `MLComputeUnitsCPUAndNeuralEngine` explicitly, and the same
> pass recorded a `powermetrics` ANE power split. The `all` story cannot be the
> whole explanation. The question is genuinely open and is tracked in
> the latency note;
> `MLComputePlan` on that exact binary settles it.

### Proving placement

| Method | What it gives you |
| --- | --- |
| `MLComputePlan` (macOS 14+) | Per-operation device assignment before running inference. Apple documents it as reporting model structure, device usage per operation, and estimated cost. |
| Instruments Core ML template | Execution timeline; an empty ANE track with busy GPU/CPU tracks means fallback. |
| `powermetrics` ANE power | The ANE domain reads `0.00 W` idle. Power off zero during inference proves engagement. |
| Espresso `[CostModelFeature]` logs | Why an op was rejected, e.g. unsupported tensor dtype. |

Activity Monitor's Energy Impact tracks CPU power and is blind to the ANE. Do
not use it as placement evidence.

```bash
sudo powermetrics --samplers cpu_power,gpu_power,ane_power -i 500
```

```bash
log stream --predicate 'subsystem == "com.apple.CoreML" || process == "ANECompilerService"' --info --debug
```

The compute-plan API shape below is from the draft and follows Apple's
documented `MLComputePlan`; check it against your SDK before relying on the
exact selector names.

```swift
let config = MLModelConfiguration()
config.computeUnits = .all

let computePlan = try await MLComputePlan.load(contentsOf: modelURL, configuration: config)
guard case let .program(program) = computePlan.modelStructure,
      let mainFunction = program.functions["main"] else { return }

for operation in mainFunction.block.operations {
    let device = computePlan.computeDeviceUsage(for: operation)
    print("\(operation.operationType) -> \(String(describing: device?.preferredComputeDevice))")
}
```

---

## Background contention

Every ANE compile goes through `ANECompilerService`, a single root-owned XPC
queue shared by the whole system. The two heaviest background clients are
`mediaanalysisd` (Photos face and scene recognition) and
`knowledgeconstructiond` (Spotlight indexing and on-device knowledge
extraction). A large batch of new files on disk sets them off, they queue
hundreds of compiles, and any application asking for the ANE waits behind them.

Killing the compiler does not help: `launchctl` respawns it immediately to
service the pending requests. Cut off the request source instead.

```bash
sudo mdutil -a -i off
```

> **Implementation evidence — measured on a dedicated M1 worker.** Copying 2.4 GB of Core
> ML artifacts onto the host started indexing;
> `knowledgeconstructiond` reached 97.6% and `ANECompilerService` 88-100%, and a
> model load requesting the ANE sat 12+ minutes at 0.0% client CPU. **Apple
> Intelligence was already disabled and made no difference** — the indexing path
> still schedules ANE work. `sudo kill -9` on the compiler produced a new PID
> back at 99.9% within seconds; killing `knowledgeconstructiond` moved the load
> to `IntelligencePlatform`. Only disabling indexing cleared it. Later,
> `mediaanalysisd` alone held 99.4% for an extended period and blocked
> measurement on its own.
>
> A useful diagnostic fell out of this: load the **same** artifact at
> `cpuOnly`, `cpuAndGPU`, and `cpuAndNeuralEngine`, then repeat with a small
> model. Measured 1.75 s / 4.42 s / hung, with a 39 MB model loading on ANE in
> 7.10 s at the same moment. Small model fine plus large model hung means an
> expensive compile; both hung would mean a genuinely stuck Neural Engine.

---

## Mitigations

### Prewarming

Load and specialize before the user asks. WhisperKit exposes this directly, and
sequential prewarming also avoids the memory spike of specializing encoder and
decoder simultaneously.

```swift
let config = WhisperKitConfig(model: "large-v3", prewarm: true)
let pipe = try await WhisperKit(config)
```

### Chunking

For models near or past the ANE compilation ceiling, the commonly reported
recipe is to split the model into several smaller sequential MLPrograms. Smaller
MIL graphs avoid the fallback-insertion blowup and produce bundles small enough
to cache reliably. *Community recipe — not verified locally.*

### Ahead-of-time compilation at install

Since bundles are hardware- and OS-specific, the compile has to happen on the
user's machine; the usual move is to pay it in the background at first launch
rather than at first inference. `MLModel.compileModel(at:)` writes to a
temporary location, so move the result somewhere durable.

```swift
let compiledModelURL = try MLModel.compileModel(at: modelDescriptionURL)

let fileManager = FileManager.default
let appSupportURL = fileManager.urls(for: .applicationSupportDirectory, in: .userDomainMask).first!
let permanentURL = appSupportURL.appendingPathComponent(compiledModelURL.lastPathComponent)
try fileManager.moveItem(at: compiledModelURL, to: permanentURL)

let model = try MLModel(contentsOf: permanentURL, configuration: config)
```

Newer SDKs offer an async variant; prefer it on a background task rather than
blocking launch.

---

## Do this / avoid this

| | Practice | Rationale |
| --- | --- | --- |
| **Do** | Verify placement with `MLComputePlan` or ANE power | Requested compute units are not evidence of where work ran |
| **Do** | Look for the cache under `~/Library/Caches/<executable-name>/` | The unqualified directory is barely used and misleads |
| **Do** | Prewarm on a background task at launch | Moves a multi-minute cold compile off the first inference |
| **Do** | Disable Spotlight indexing on dedicated inference hosts | It is a first-order ANE contender, independent of Apple Intelligence |
| **Avoid** | Profiling ANE behaviour under `MLComputeUnits.all` | Documented as an OS choice; it can schedule around the ANE silently |
| **Avoid** | Treating a fast load as proof of ANE residency | A fast load may mean the ANE compile never happened |
| **Avoid** | Inferring compile progress from cache size | The bundle is written only at the end, and probably not where you are looking |
| **Avoid** | Clearing `~/Library/Caches` on an inference host | Discards gigabytes of ANE bundles and forces cold recompiles |

---

## Sources

The raw report's numbered citations resolve to Google grounding-redirect URLs
that are not stable references. The substantive external claims trace to Apple's
Core ML documentation (`MLComputeUnits`, `MLComputePlan`,
`MLModel.compileModel(at:)`), Apple's *Deploying Transformers on the Apple
Neural Engine* article and the `ane-transformers` repository, WhisperKit's
prewarming configuration, and Apple Developer Forums / GitHub issue threads on
`ANECompilerService`, `e5bundlecache`, and slow first model loads. Consult the
