MissingManual

ANE / cold start · Collective Library

ANE Compile Cost and Compiled-Bundle Caching for Large Core ML Models

Collective Library edition. This is the complete technical report. Private filesystem paths, internal run identifiers, campaign-control notes, and repository navigation were removed. Technical claims, code, measurements, evidence labels, citations, corrections, and falsification criteria are preserved.

Why a large Core ML encoder can take twenty minutes to load the first time, where the compiled artifact goes afterwards, and how to prove the Neural Engine was actually used.

Provenance. Ingested 2026-08-07 from a Gemini Deep Research Max run:

  apple-neural-engine-compilation-cost-and-compiled-bundle-caching-for-large-core-/
  2026-08-08T03-52-03-618Z/raw-report.md

SHA-256 f3b30252b62ac1f06f139735234e267fc3ae646d3394a0e3725689bf76aa6695

Compute-unit and MLComputePlan semantics re-checked against Apple's Core ML documentation via Context7. Sections marked Implementation evidence are firsthand measurements from the reference implementation and take precedence over the draft where they conflict.

Executive summary

  • A cold ANE compile of a ~1 GB transformer encoder takes 10-50 minutes on M1, and practitioners report this as normal rather than pathological. Cost is driven by graph segmentation and fallback insertion, not parameter count.
  • The compiled bundle is cached per calling executable, under ~/Library/Caches/<executable-name>/com.apple.e5rt.e5bundlecache/. The cache key is a digest over the MIL program, bound to the OS build and the caller's cache domain.
  • MLComputeUnits.all is not a directive to use the Neural Engine. Apple documents it as letting the OS choose; it can load fast by scheduling around the ANE entirely. Only MLComputePlan, an Instruments Core ML trace, or ANE power in powermetrics proves placement.
  • ANECompilerService is a single shared root-owned queue. Spotlight indexing and Photos media analysis saturate it, and any ANE model load then waits behind them.
  • The ANE has no general matrix-multiply unit. It executes matmuls as 1x1 convolutions, which is why einsum-heavy attention graphs compile slowly and why Apple's ane-transformers rewrites attention into 4D Conv2d.

Compiled-bundle caching

Location and cache key

Compiled ANE artifacts ("E5 binaries", after Apple's Execution Engine 5 runtime) are written to ~/Library/Caches/<identifier>/com.apple.e5rt.e5bundlecache/. Below that the layout is <os-build>/<hash>/, where the hash is a digest over the MIL program. The cache key therefore includes the model graph and weights, the exact OS build, and the caller's cache domain.

System daemons get their own domains — media analysis caches under its own container, Spotlight under com.apple.Spotlight.

Because the key includes the OS build and because ANE topology differs across Apple Silicon generations, a bundle cannot be pre-built by a developer and shipped. The compile happens on the end user's machine.

Implementation evidence — the identifier is the executable name, and the shared directory is a red herring. Measured on an M1 Mac mini running macOS 15.7.7:

~/Library/Caches/cu/com.apple.e5rt.e5bundlecache               4863 MB
~/Library/Caches/encoder-bench/com.apple.e5rt.e5bundlecache    3646 MB
~/Library/Caches/kokoro-worker/com.apple.e5rt.e5bundlecache    9014 MB
~/Library/Caches/whisper-cli/com.apple.e5rt.e5bundlecache         1 MB
~/Library/Caches/com.apple.e5rt.e5bundlecache                    29 MB

cu and encoder-bench are bare command-line binaries with no bundle identifier, and each got a cache domain named after the executable. The unqualified com.apple.e5rt.e5bundlecache at the top of Caches is barely used. Rename or rebuild a binary and it starts from a cold cache, because the domain follows the executable name.

Size ceilings and cache failure

The draft attributes multi-minute recompiles on every load to an undocumented ~2 GB ceiling, citing practitioner reports of BNNS Graph Compile: failed to preallocate file with error: No space left on device when a compiled bundle grows past roughly 2 GB, with jetsam killing the process or the write being abandoned. Reports of that error are real and worth knowing.

Implementation evidence — the "silently fails to persist" story did not hold here, and believing it cost hours. A 1.2 GB Whisper large-v3-turbo encoder compiled for cpuAndNeuralEngine in 985-1300 s cold, and the identical load with the bundle cached took 0.64 s. Caching worked. The apparent "nothing persists" symptom came entirely from watching the wrong directory: the unqualified com.apple.e5rt.e5bundlecache never grew because the compiler was writing to ~/Library/Caches/cu/… and ~/Library/Caches/encoder-bench/… instead. Two published conclusions were retracted over this. Check the per-executable domain before concluding a persist failure.

Eviction

Cached bundles are volatile. OS updates, low disk space, and long disuse purge them, so a load that was fast last month can be slow today. Applications that care about first-inference latency prewarm rather than assume residency.

Implementation evidence — treat a cache directory as load-bearing infrastructure. A speech worker on the same host held 9 GB of ANE bundles. Clearing ~/Library/Caches on that machine, or shipping a renamed worker binary, would discard all of it and force a full cold recompile on the next model load.


What makes ANE compilation slow

Graph segmentation and fallback insertion

The compiler tries to place every operation on the Neural Engine. When it hits one the ANE cannot execute — because of precision, dynamic control flow, or tensors exceeding ANE SRAM — it must cut the graph and insert a transfer to CPU or GPU. On large graphs this produces hundreds of "islands"; one reported case inserted 526 ANE-to-CPU/GPU context transfers and 262 single-op fallback segments. Compilation passes then do work proportional to layers times transfers, which is where tens of minutes come from.

Compile cost tracks graph structure, not parameter count. A model that maps cleanly compiles far faster than a smaller one that does not.

einsum versus Conv2d

The ANE has no general matrix-multiply unit. It executes matmuls through its convolution datapath as 1x1 convolutions. einsum is valid Core ML MIL, but an attention graph built from many einsum ops has to be decomposed into convolution tiles or handed to CPU fallbacks for layout reasons.

This is why Apple's ane-transformers reference rewrites attention into 4D Conv2d form and removes einsum entirely. The ANE also wants strict stride alignment — the last axis a multiple of 64 bytes — and unaligned tensors force the compiler to insert layout-conversion layers.

Implementation evidence — a stock Whisper encoder export is einsum-heavy. The compiled large-v3-turbo encoder used here reports Einsum 1280, Softmax 640, Conv 194, LayerNorm 65 in its MIL op histogram, and the same histogram appears in the artifact produced by whisper.cpp's own converter. Both took 985-1300 s to compile cold, so the long compile is a property of this model shape rather than of any one exporter. Whether the einsum count is causal here was not tested; that would need MLComputePlan op-level placement, which is the check to run before acting on this hypothesis.


Compute units and silent fallback

Apple documents the cases plainly: all lets the OS choose the best option across CPU, GPU, and Neural Engine and is the default; cpuOnly, cpuAndGPU, and cpuAndNeuralEngine restrict it (verified via Context7 against Apple's MLComputeUnits documentation).

The practical consequence the draft draws out: all is permission, not a requirement. If ANE compilation is expensive or the graph maps poorly, the planner can schedule the model onto GPU/CPU and return quickly, while cpuAndNeuralEngine forces the expensive ANE path. A fast load under all and a twenty-minute load under cpuAndNeuralEngine for the same artifact are consistent with this.

Implementation evidence — the fast-load-means-all explanation does not fit one important local case, so treat it as a hypothesis. The draft proposes that a previously recorded 465.9 ms load in the reference implementation must have been all quietly choosing the GPU. But the build that produced that measurement was patched to request MLComputeUnitsCPUAndNeuralEngine explicitly, and the same pass recorded a powermetrics ANE power split. The all story cannot be the whole explanation. The question is genuinely open and is tracked in the latency note; MLComputePlan on that exact binary settles it.

Proving placement

Method What it gives you
MLComputePlan (macOS 14+) Per-operation device assignment before running inference. Apple documents it as reporting model structure, device usage per operation, and estimated cost.
Instruments Core ML template Execution timeline; an empty ANE track with busy GPU/CPU tracks means fallback.
powermetrics ANE power The ANE domain reads 0.00 W idle. Power off zero during inference proves engagement.
Espresso [CostModelFeature] logs Why an op was rejected, e.g. unsupported tensor dtype.

Activity Monitor's Energy Impact tracks CPU power and is blind to the ANE. Do not use it as placement evidence.

sudo powermetrics --samplers cpu_power,gpu_power,ane_power -i 500
log stream --predicate 'subsystem == "com.apple.CoreML" || process == "ANECompilerService"' --info --debug

The compute-plan API shape below is from the draft and follows Apple's documented MLComputePlan; check it against your SDK before relying on the exact selector names.

let config = MLModelConfiguration()
config.computeUnits = .all

let computePlan = try await MLComputePlan.load(contentsOf: modelURL, configuration: config)
guard case let .program(program) = computePlan.modelStructure,
      let mainFunction = program.functions["main"] else { return }

for operation in mainFunction.block.operations {
    let device = computePlan.computeDeviceUsage(for: operation)
    print("\(operation.operationType) -> \(String(describing: device?.preferredComputeDevice))")
}

Background contention

Every ANE compile goes through ANECompilerService, a single root-owned XPC queue shared by the whole system. The two heaviest background clients are mediaanalysisd (Photos face and scene recognition) and knowledgeconstructiond (Spotlight indexing and on-device knowledge extraction). A large batch of new files on disk sets them off, they queue hundreds of compiles, and any application asking for the ANE waits behind them.

Killing the compiler does not help: launchctl respawns it immediately to service the pending requests. Cut off the request source instead.

sudo mdutil -a -i off

Implementation evidence — measured on a dedicated M1 worker. Copying 2.4 GB of Core ML artifacts onto the host started indexing; knowledgeconstructiond reached 97.6% and ANECompilerService 88-100%, and a model load requesting the ANE sat 12+ minutes at 0.0% client CPU. Apple Intelligence was already disabled and made no difference — the indexing path still schedules ANE work. sudo kill -9 on the compiler produced a new PID back at 99.9% within seconds; killing knowledgeconstructiond moved the load to IntelligencePlatform. Only disabling indexing cleared it. Later, mediaanalysisd alone held 99.4% for an extended period and blocked measurement on its own.

A useful diagnostic fell out of this: load the same artifact at cpuOnly, cpuAndGPU, and cpuAndNeuralEngine, then repeat with a small model. Measured 1.75 s / 4.42 s / hung, with a 39 MB model loading on ANE in 7.10 s at the same moment. Small model fine plus large model hung means an expensive compile; both hung would mean a genuinely stuck Neural Engine.


Mitigations

Prewarming

Load and specialize before the user asks. WhisperKit exposes this directly, and sequential prewarming also avoids the memory spike of specializing encoder and decoder simultaneously.

let config = WhisperKitConfig(model: "large-v3", prewarm: true)
let pipe = try await WhisperKit(config)

Chunking

For models near or past the ANE compilation ceiling, the commonly reported recipe is to split the model into several smaller sequential MLPrograms. Smaller MIL graphs avoid the fallback-insertion blowup and produce bundles small enough to cache reliably. Community recipe — not verified locally.

Ahead-of-time compilation at install

Since bundles are hardware- and OS-specific, the compile has to happen on the user's machine; the usual move is to pay it in the background at first launch rather than at first inference. MLModel.compileModel(at:) writes to a temporary location, so move the result somewhere durable.

let compiledModelURL = try MLModel.compileModel(at: modelDescriptionURL)

let fileManager = FileManager.default
let appSupportURL = fileManager.urls(for: .applicationSupportDirectory, in: .userDomainMask).first!
let permanentURL = appSupportURL.appendingPathComponent(compiledModelURL.lastPathComponent)
try fileManager.moveItem(at: compiledModelURL, to: permanentURL)

let model = try MLModel(contentsOf: permanentURL, configuration: config)

Newer SDKs offer an async variant; prefer it on a background task rather than blocking launch.


Do this / avoid this

Practice Rationale
Do Verify placement with MLComputePlan or ANE power Requested compute units are not evidence of where work ran
Do Look for the cache under ~/Library/Caches/<executable-name>/ The unqualified directory is barely used and misleads
Do Prewarm on a background task at launch Moves a multi-minute cold compile off the first inference
Do Disable Spotlight indexing on dedicated inference hosts It is a first-order ANE contender, independent of Apple Intelligence
Avoid Profiling ANE behaviour under MLComputeUnits.all Documented as an OS choice; it can schedule around the ANE silently
Avoid Treating a fast load as proof of ANE residency A fast load may mean the ANE compile never happened
Avoid Inferring compile progress from cache size The bundle is written only at the end, and probably not where you are looking
Avoid Clearing ~/Library/Caches on an inference host Discards gigabytes of ANE bundles and forces cold recompiles

Sources

The raw report's numbered citations resolve to Google grounding-redirect URLs that are not stable references. The substantive external claims trace to Apple's Core ML documentation (MLComputeUnits, MLComputePlan, MLModel.compileModel(at:)), Apple's Deploying Transformers on the Apple Neural Engine article and the ane-transformers repository, WhisperKit's prewarming configuration, and Apple Developer Forums / GitHub issue threads on ANECompilerService, e5bundlecache, and slow first model loads. Consult the