# Physical whisper.cpp + Core ML Encoder Verification on Apple Silicon (2026)

> **Collective Library edition.** This is the complete technical report. Private filesystem paths, internal run identifiers, campaign-control notes, and repository navigation were removed. Technical claims, code, measurements, evidence labels, citations, corrections, and falsification criteria are preserved.

August 1, 2026

This guide defines how to *prove*, from a physical run, that a pinned
whisper.cpp build actually executed a Core ML encoder instead of its ordinary
GGML encoder. It exists because the Phase 7 release gate treats a silent
ordinary-encoder run as a FAILED gate, and because a passing transcript is not
evidence of which encoder produced it.

Scope: the runtime boundary only — build flags, artifact naming, the exact log
contract, the headless subprocess contract, parity binding, and compute-device
observation. Model conversion theory lives in the Core ML guides; the container
question lives in the runtime-contract guide.

**Verified against:**

| Source | Revision / version | Checked |
| --- | --- | --- |
| whisper.cpp (repo pin) | `97c56f1dc1d1100a9d859c865a20c82d22f823ed` plus patch `99bf1e4a9f768745bbaad61d4d131e5360c2f240e2fa577ce64726d9c67bfcef` | 2026-08-01 |
| whisper.cpp `master` | `2ca53bb45e38748d07b310eeb36245a7157ac882` | 2026-08-01 |
| coremltools | `9.0` (released 2025-11-10; current latest on PyPI) | 2026-07-30 |
| Hugging Face Transformers | `main` (`modeling_whisper.py`) | 2026-07-30 |
| Host used for local tool checks | macOS 15.7.7 (24G720), Xcode installed | 2026-07-30 |

At both checked source revisions, `whisper_state.batch` is still declared
without initialization, the required Core ML failure returns before
`whisper_batch_init`, and teardown unconditionally calls `whisper_batch_free`.
Do not infer whole-file identity from that lifecycle match; line numbers below
name the pinned source unless stated otherwise.

This guide was refreshed from the Deep Research Max draft at
The draft's claim that issue #3807 documents this defect was removed: #3807 is
about unvalidated zero-dimension model parameters, not `whisper_batch` teardown.
The lifecycle claims retained here were checked against the pinned and current
upstream source plus the implementationsitory's repeated physical A/B.

## 1. Direct verdict

| Question | Verified answer |
| --- | --- |
| Which single log line proves the Core ML encoder loaded? | `whisper_init_state: Core ML model loaded` |
| Does that line also prove the GGML encoder was skipped? | Yes. The same pointer that gates the line gates `whisper_encode_external()`. See section 2.2. |
| Does it prove the Apple Neural Engine ran? | No. Nothing in whisper.cpp observes device placement. |
| Where do the log lines go? | **stderr**, always. Never stdout. |
| Can the gate lose the log lines? | Yes — `whisper-cli -np/--no-prints` installs a null log callback and silences them. |
| Is silent fallback possible by default? | Not from a *load failure*: `WHISPER_COREML_ALLOW_FALLBACK` defaults `OFF`, so init hard-fails. It is possible from a *build* that simply lacks `WHISPER_COREML=1`. |
| What does the build-level marker look like? | `COREML = 1` inside the `system_info:` line. `COREML = 0` (or an absent marker) means the binary has no Core ML code at all. |
| Which encoder input shape does the runtime hand to Core ML? | `[1, n_mels, 2 * n_audio_ctx]` = `[1, 80, 3000]` for standard models. |
| Which compute units does whisper.cpp request? | `MLComputeUnitsAll`, hardcoded. Not configurable without patching source. |
| Is `.mlpackage` parity the same evidence as `.mlmodelc` parity? | No. The runtime loads the compiled `.mlmodelc`; bind parity to those exact bytes. |
| What identifies the supported downstream runtime? | Upstream commit + exact compatibility-patch SHA-256 + built `whisper-cli` SHA-256. |
| May the negative control crash? | No. It must exit exactly `3` and emit both the Core ML load-failure and final CLI initialization-failure markers. |

## 2. The load-confirmation contract

### 2.1. Every Core ML line the runtime can emit

All four come from `whisper_init_state()`, so `__func__` expands to
`whisper_init_state`. Each row's text is the corresponding source format string
with `%s` substituted — the literal portions are exact, and the only variable
part is the path:

| Emitted text (after `%s` expansion) | Source | Meaning |
| --- | --- | --- |
| `whisper_init_state: loading Core ML model from '<path>'` | `src/whisper.cpp:3443` | Core ML code is compiled in and a path was derived. Proves nothing about success. |
| `whisper_init_state: first run on a device may take a while ...` | `src/whisper.cpp:3444` | Always printed, unconditionally, before the load attempt. Not a latency measurement. |
| `whisper_init_state: failed to load Core ML model from '<path>'` | `src/whisper.cpp:3448` | Load failed. Logged at ERROR level. |
| `whisper_init_state: Core ML model loaded` | `src/whisper.cpp:3454` | **The load-confirmation line.** Emitted only on the `else` branch of `if (!state->ctx_coreml)`. |

Build-level marker, from `whisper_print_system_info()`:

```text
s += "COREML = "    + std::to_string(whisper_has_coreml())     + " | ";
```

`src/whisper.cpp:4334`, where `whisper_has_coreml()` (`src/whisper.cpp:4313`)
returns `1` only under `#ifdef WHISPER_USE_COREML`. `whisper-cli` embeds the
whole string in one `system_info:` line (`examples/cli/cli.cpp:1187`–`1188`).

The line's composition, derived from those two sources rather than quoted from
a run: `system_info: n_threads = <n> / <m> | ` then `WHISPER : COREML = <0|1> |
OPENVINO = <0|1> | ` then one segment per registered ggml backend with its
feature flags. The backend tail varies by machine and by build, and the
upstream README's example still shows a pre-backend-registry format. Match on
the substring `COREML = 1`. Never match the whole line.

### 2.2. Why the load line is sufficient proof of encoder substitution

This is the chain the gate rests on, and it is airtight in current source:

1. `state->ctx_coreml` is assigned at `src/whisper.cpp:3446`.
2. `whisper_init_state: Core ML model loaded` is printed **iff**
   `state->ctx_coreml != nullptr` (`src/whisper.cpp:3447`–`3454`).
3. `whisper_encode_external()` returns `use_coreml`, defined as
   `wstate.ctx_coreml != nullptr` (`src/whisper.cpp:1964`), i.e. exactly the
   same predicate.
4. In `whisper_encode_internal()`, when `whisper_encode_external()` is true the
   conv graph compute is skipped and `whisper_coreml_encode(...)` runs instead
   (`src/whisper.cpp:2405`–`2413`); the encoder graph block is then also
   skipped (`src/whisper.cpp:2421`).
5. In `whisper_build_graph_conv()`, the external branch does not even build the
   convolution ops — it allocates a bare `embd_enc` input tensor "the external
   encoder will write into" (`src/whisper.cpp:2020`–`2026`).

So the load marker is not a hint. It is the same boolean that removes the GGML
encoder from the graph.

### 2.3. The assertion the Phase 7 gate must make

Four conditions, all required, from one physical run:

1. exit status `0`;
2. `COREML = 1` present (build actually has Core ML compiled in);
3. `whisper_init_state: Core ML model loaded` present;
4. `whisper_init_state: failed to load Core ML model` **absent**.

Condition 4 matters only if a build ever sets `WHISPER_COREML_ALLOW_FALLBACK`,
but asserting it costs nothing and converts a misconfigured build from a silent
pass into a loud failure.

the implementationsitory encodes conditions 2–4 as
`COREML_ENABLED_MARKER`, `COREML_LOAD_MARKER`, and
`COREML_FALLBACK_FAILURE_MARKER`. The shared failure-marker and termination
contract lives in `whisper_tuner/export/runtime_contracts.py`; the physical
probe lives in `whisper_tuner/export/coreml_fallback_probe.py`.

### 2.4. The negative control

Assertions on a passing run can be satisfied by a stale log. Pair every gate
with a *fail-closed canary*: copy only the `ggml-*.bin` into a scratch
directory — deliberately leaving the `.mlmodelc` behind — and run the same
binary. Required outcome:

- exit exactly `3`;
- `failed to load Core ML model` present;
- `error: failed to initialize whisper context` present;
- no retry and no signal termination (`-6`, `134`, or otherwise).

Chain, verified in source: `whisper_init_state` returns `nullptr` when Core ML
fails and fallback is not enabled (`src/whisper.cpp:3449`–`3452`) →
`whisper_init_from_file_with_params` returns `nullptr`
(`src/whisper.cpp:3755`–`3759`) → `whisper-cli` prints
`error: failed to initialize whisper context` and returns `3`
(`examples/cli/cli.cpp:1083`–`1086`).

If that canary succeeds, crashes, times out, or lacks either marker, the build
has not proved a clean fail-closed refusal and no positive result from it is
admissible.

## 3. Pinning and building the runtime

```bash
git clone https://github.com/ggml-org/whisper.cpp.git
cd whisper.cpp
git checkout 97c56f1dc1d1100a9d859c865a20c82d22f823ed
git apply /path/to/the/exact-compatibility.patch
git diff --no-ext-diff --binary -- src/whisper.cpp
# output must byte-match the declared patch; status must be exactly:
#  M src/whisper.cpp

cmake -B build-coreml \
  -DWHISPER_COREML=1 \
  -DWHISPER_COREML_ALLOW_FALLBACK=OFF
cmake --build build-coreml -j --config Release
```

Both options default `OFF` (`CMakeLists.txt:92`–`93`). `-DWHISPER_COREML=1`
locates the `CoreML` and `Foundation` frameworks, adds `-DWHISPER_USE_COREML`,
and compiles the `whisper.coreml` target with `-fobjc-arc`
(`src/CMakeLists.txt:30`–`83`). A missing framework is a hard
`FATAL_ERROR`, not a downgrade (`src/CMakeLists.txt:39`).

Build-hygiene rules that have actually bitten this class of gate:

- Use a **separate build directory** from the non-Core-ML build. A stale
  `build/` containing a non-Core-ML `whisper-cli` will pass every transcript
  test and fail every load assertion — or worse, be silently reused.
- Record `sha256` of the produced `whisper-cli`, not just the source SHA.
- Re-verify immediately before and after the build that the checkout contains
  exactly the declared one-line patch and no other tracked or untracked change.
  A generically "clean" checkout is unpatched and unsupported; a generically
  dirty checkout is unverifiable.

## 4. The artifact contract

### 4.1. Filename derivation is mechanical and unforgiving

`whisper_get_coreml_path_encoder()` (`src/whisper.cpp:3328`–`3345`) transforms
the **exact `-m` argument string**:

1. truncate at the last `.` in the whole path;
2. if the trailing `-` segment is exactly five characters and matches
   `-q?_?`, drop it;
3. append `-encoder.mlmodelc`.

| `-m` argument | Derived Core ML path | Note |
| --- | --- | --- |
| `models/ggml-base.en.bin` | `models/ggml-base.en-encoder.mlmodelc` | Upstream convention. |
| `.../ggml-model.bin` | `.../ggml-model-encoder.mlmodelc` | the reference implementation's exported name. |
| `.../ggml-base.en-q5_0.bin` | `.../ggml-base.en-encoder.mlmodelc` | Quantized models share the FP encoder. |
| `.../ggml-model-q8_0.bin` | `.../ggml-model-encoder.mlmodelc` | Same rule. |
| `.../ggml-tune-q1_5.bin` | `.../ggml-tune-encoder.mlmodelc` | Rule is **syntactic**, not a quant-mode allowlist. Any `-q?_?` suffix is eaten. |
| `/data/v1.4/ggml-model` (no extension) | `/data/v1-encoder.mlmodelc` | Step 1 finds the dot in the *directory*. Always give the model a `.bin` extension. |

Because derivation uses the argument string, the `.mlmodelc` must live beside
the model **as the caller spelled it**. Copying the `.bin` to a temp directory
without the `.mlmodelc` is the negative control of section 2.4 — and the
accidental version of it is a common false failure.

### 4.2. Tensor contract

| Boundary | Value | Source |
| --- | --- | --- |
| Core ML input feature name | `logmel_data` | `src/coreml/whisper-encoder-impl.m:25`, `:29` |
| Core ML output feature name | `output` | `src/coreml/whisper-encoder-impl.m:180` |
| Input `MLMultiArray` shape | `[1, n_mel, n_ctx]` where `n_ctx` is the ggml mel width | `src/coreml/whisper-encoder.mm:55`–`62` |
| ggml mel tensor | `ggml_new_tensor_2d(F32, 2*n_audio_ctx, n_mels)` → `3000 x 80` | `src/whisper.cpp:1997` |
| Effective input | `[1, 80, 3000]` (`[1, 128, 3000]` for large-v3 class) | derived from the two rows above |
| Input dtype | `MLMultiArrayDataTypeFloat32` | `src/coreml/whisper-encoder.mm:58` |
| Expected output element count | `n_audio_ctx * n_audio_state` (e.g. `1500 * 384 = 576,000` for tiny) | `src/whisper.cpp:2022` |
| Output copy | `memcpy(out, ...dataPointer, count * sizeof(float))` | `src/coreml/whisper-encoder.mm:67` |
| Compute units requested | `MLComputeUnitsAll` (hardcoded) | `src/coreml/whisper-encoder.mm:29` |

Three consequences follow directly from that `memcpy` and are worth stating
plainly, because none of them produce an error message:

- **Float32 output is mandatory.** The copy multiplies the element count by
  `sizeof(float)`. A Core ML model whose *interface* dtype is Float16 would be
  read past its buffer. (Upstream's `--quantize` sets
  `compute_precision=FLOAT16` but leaves the IO type unspecified, so the
  interface stays fp32; a hand-rolled fp16-IO export would not.)
- **An oversized output overruns `embd_enc`.** There is no shape check on the
  runtime side. Validate input and output shapes on the Python side before the
  artifact is ever placed beside the model.
- **A failed prediction is silent.** Both `initWithContentsOfURL:` and
  `predictionFromLogmel_data:` are called with `error:nil`
  (`src/coreml/whisper-encoder.mm:31`, `:65`). A prediction failure yields a
  nil output, `count` of `0`, a zero-length copy, and an encoder state that was
  never written — producing empty or hallucinated text with a clean log and a
  `Core ML model loaded` line. **This is the failure mode the load marker
  cannot catch.** Only transcript parity catches it.

### 4.3. Compilation to `.mlmodelc` requires full Xcode

```bash
xcrun coremlc compile path/to/ggml-model-encoder.mlpackage path/to/output-dir/
# produces path/to/output-dir/ggml-model-encoder.mlmodelc
```

Locally verified on macOS 15.7.7: `xcrun --find coremlc` resolves to
`/Applications/Xcode.app/Contents/Developer/Toolchains/XcodeDefault.xctoolchain/usr/bin/coremlc`.
`/Library/Developer/CommandLineTools` was installed on the same host and
contains **no** `coremlc`. Command Line Tools alone are not sufficient; a CI
runner with only CLT will fail here, not at conversion.

`coremlc` names the output after the input basename, which is why the
`.mlpackage` must already be named `ggml-<basename>-encoder.mlpackage`. This is
the same mechanism upstream relies on before renaming
(`models/generate-coreml-model.sh:31`–`33`).

## 5. Producing the encoder: upstream script versus repo-owned converter

### 5.1. Current state of `generate-coreml-model.sh`

`models/generate-coreml-model.sh` (38 lines) calls
`models/convert-whisper-to-coreml.py --encoder-only True --optimize-ane True`
for named OpenAI models, or `convert-h5-to-coreml.py` under `-h5` for a Hugging
Face directory, then compiles and renames.

Facts about the current converter, from source:

| Property | Current value | Source |
| --- | --- | --- |
| Input shape | `(1, hparams.n_mels, 3000)` — **fixed**, no `RangeDim`, no `EnumeratedShapes` | `models/convert-whisper-to-coreml.py:260` |
| Backend | `convert_to="mlprogram"` | `:266` |
| Precision | `FLOAT32` unless `--quantize` (which the shell script never passes) | `:270` |
| Compute units at convert time | `ct.ComputeUnit.ALL` | `:269` |
| SDPA | globally disabled via `whisper.model.MultiHeadAttention.use_sdpa = False` | `:31` |
| ANE reshaping | `nn.Linear` → 1x1 `nn.Conv2d`, `LayerNormANE`, 4D `(B, C, 1, S)` | `:34`–`:166` |
| Model allowlist | fixed list of OpenAI names; arbitrary names rejected | `:307` |

### 5.2. Corrections to the historical guide

Three claims in `coreml-hybrid-whispercpp.md` do not match current source:

| Historical claim | Current reality |
| --- | --- |
| "Uses `ct.RangeDim(1, 3000)` for variable audio lengths" | The encoder input is a **fixed** `(1, n_mels, 3000)` tensor. There is no flexible shape anywhere in the current converter. |
| "`make WHISPER_COREML=1`" | The Makefile path is gone; CMake with `-DWHISPER_COREML=1` is the documented build. |
| "`np.allclose(torch_out, coreml_out, atol=1e-2)  # FP16 tolerance`" | Far too loose for a release gate, and max-absolute-error is the wrong statistic entirely. See section 7. |

### 5.3. The dependency trap in the upstream path

`ane_transformers` is required by `convert-whisper-to-coreml.py:10`. Its
`install_requires` is:

```python
"torch>=1.10.0,<=1.11.0",
"coremltools>=5.2.0",
"transformers>=4.18.0",
"protobuf>=3.1.0,<=3.20.1",
```

It is distributed as an **sdist only** and its last release, 0.1.3, was
uploaded 2022-08-09 — no release in roughly four years. Installing it into this
repository's environment (`torch==2.11.0`, `coremltools==9.0`) is a resolver
conflict or a catastrophic downgrade. If the upstream reference conversion is
ever needed for comparison, build it in a **throwaway isolated environment**;
never in the project environment.

Independently, `coremltools==9.0` declares `_TORCH_MAX_VERSION = "2.7.0"`
(`coremltools/_deps/__init__.py:158`) and only *warns* when exceeded:

```text
Torch version 2.11.0 has not been tested with coremltools. You may run into
unexpected errors. Torch 2.7.0 is the most recent version that has been tested.
```

`_warn_if_above_max_supported_version` uses `logger.warning`, not an exception
(`coremltools/_deps/__init__.py:33`–`39`). Treat that warning as a recorded
build fact, not noise — it is the single most likely explanation offered for
unexplained numeric drift, and the implementationsitory has already tested it (Torch 2.7
did not change the FP32 outcome; see the debug notes).

The related `_SKLEARN_MAX_VERSION = "1.5.1"`
(`coremltools/_deps/__init__.py:57`) is the reason the reference implementation caps
scikit-learn: above it, coremltools disables its sklearn converter at import.

### 5.4. Why the reference implementation converts from Hugging Face directly

The upstream script loads `openai-whisper` checkpoints by name. A fine-tuned
Hugging Face artifact is not on that allowlist, and round-tripping through
`convert-h5-to-coreml.py` adds an unpinned second converter. the implementationsitory's
`whisper_tuner/export/coreml.py` instead traces the HF encoder directly and
must therefore reproduce the runtime contract itself:

- input named `logmel_data`, fixed `(1, n_mels, 3000)`, fp32;
- output named `output`, `(1, 1500, hidden_size)`, fp32;
- `mlprogram`;
- compiled to `ggml-model-encoder.mlmodelc` beside `ggml-model.bin`.

The 3000-versus-1500 confusion is worth naming because it is the natural error
here: `config.max_source_positions` is **1500** (the encoder's *output*
length), while the required mel width is **3000**. Hugging Face enforces this
and produces an exact error:

```python
expected_seq_length = self.config.max_source_positions * self.conv1.stride[0] * self.conv2.stride[0]
if input_features.shape[-1] != expected_seq_length:
    raise ValueError(
        f"Whisper expects the mel input features to be of length {expected_seq_length}, "
        f"but found {input_features.shape[-1]}. Make sure to pad the input mel features to {expected_seq_length}."
    )
```

(`transformers/models/whisper/modeling_whisper.py:612`–`616`.) `1500 * 1 * 2 =
3000`. Derive the contract from `max_source_positions` and the two conv
strides; never hardcode either number.

## 6. Driving the pinned runtime headlessly from Python

### 6.1. Stream and flag contract

- **All library logs go to stderr.** The default log callback is
  `fputs(text, stderr); fflush(stderr);` (`src/whisper.cpp:9186`–`9187`).
  `whisper-cli`'s own diagnostics also use `fprintf(stderr, ...)`. Capture both
  streams and search the concatenation; do not search stdout alone.
- **Never pass `-np` / `--no-prints`.** It installs `cb_log_disable`, an empty
  log callback (`examples/cli/cli.cpp:1047`–`1048`), which erases every Core ML
  line. A gate run with `-np` cannot pass and cannot fail honestly — it is
  simply blind.
- **Never pass `-ac` / `--audio-ctx`** (`examples/cli/cli.cpp:168`). It changes
  the ggml mel width to `2 * audio_ctx`, which no longer matches the encoder's
  fixed 3000-frame input. No guard exists in source; the result is the silent
  nil-prediction path of section 4.2.
- **Keep `-p` at its default of 1.** `whisper_full_parallel` calls
  `whisper_init_state(ctx)` once per extra processor
  (`src/whisper.cpp:7831`–`7833`), so `-p N` performs N Core ML loads, emits N
  load markers, and holds N model instances. Fine for throughput, hostile to a
  deterministic gate.
- **Pin decoding.** `--language <lang> --no-timestamps --temperature 0
  --temperature-inc 0` removes temperature-fallback nondeterminism from the
  transcript comparison.
- **Sanitize the child environment.** Third-party subprocesses should not
  inherit Hugging Face or cloud credentials.

### 6.2. Exit codes

Verified from `examples/cli/cli.cpp`:

| Code | Meaning | Line |
| --- | --- | --- |
| `0` | Success | `1355` |
| `1` | Argument parse failure, or missing `@response` file | `994`, `1013` |
| `2` | No input file specified | `1032` |
| `3` | Unknown DTW preset, **or `whisper_init_from_file_with_params` returned null** | `1077`, `1085` |
| `4` | Grammar parse failure | `1106` |
| `10` | `whisper_full_parallel` failed to process audio | `1320` |

`3` is the Core-ML-relevant one. Do not treat "nonzero" as sufficient for the
negative control: require `3` **and** the failure marker, otherwise a typo in
the model path produces the same "expected failure".

#### Why unpatched exit 3 is intermittent — and never acceptable

The source intends to return `3`, but the pinned upstream tree leaves
`whisper_state.batch` uninitialized:

```cpp
whisper_state * state = new whisper_state;
// ... backend, KV-cache, alignment-mask, and Core ML setup ...
whisper_batch batch;  // member declaration, no initializer
// Core ML required-load failure calls whisper_free_state(state) here
state->batch = whisper_batch_init(...);  // not reached
```

`whisper_free_state` then unconditionally calls
`whisper_batch_free(state->batch)`. That function conditionally frees five
pointer members, but the conditions themselves read indeterminate pointers.
The resulting invalid free explains the intermittent `SIGABRT`; the earlier
Metal-race explanation was wrong.

The supported patch is deliberately one line:

```diff
-    whisper_batch batch;
+    whisper_batch batch = {};
```

Aggregate zero-initialization makes each pointer null before any early return.
The normal path later overwrites the member with `whisper_batch_init`, so the
successful initialization order is unchanged.

Measured A/B on August 1, 2026:

| Source | Unpatched | Exact zero-init patch |
| --- | ---: | ---: |
| pinned `97c56f1…` | 22/30 exit `3`; 8/30 `SIGABRT` | 100/100 exit `3` |
| then-current upstream master | 23/30 exit `3`; 7/30 `SIGABRT` | 100/100 exit `3` |

Issue #2750 independently records the same Core ML failure endpoint and invalid
free, but it does not establish the reference implementation's root-cause analysis. Issue
#3807 is unrelated and must not be cited for this defect.

The release gate accepts only the clean, fully observed path: exact exit `3`,
the Core ML load-failure marker, and the final CLI initialization-failure
marker. A crash proves memory unsafety, not fallback discipline. It is a hard
failure even if a marker was printed. Retrying is forbidden because it would
turn undefined behavior into probabilistic success.

The exporter and receipt validator share the classifier and intrinsic evidence
predicate from `whisper_tuner.export.runtime_contracts`; invalid exits, missing
markers, malformed evidence hashes, and relabeled crashes all fail closed.

### 6.3. Timeouts

The first positive run against a freshly compiled `.mlmodelc` triggers on-device ANE
compilation. The upstream README states this directly: the first run is slow
because the ANE service compiles the model to a device-specific format, and
later runs are faster.

Practical consequence for an automated gate:

- Give the **first** invocation a generous timeout (minutes, not seconds).
  Community reports of 20–50 s first loads for large encoders are common and
  plausible; treat any single published number as an estimate, not a bound.
- A **positive-path warm-up invocation** is allowed before the measured positive
  run; record both wall-clock values rather than gating on either. Never warm or
  retry the negative control—the first canary result is the evidence.
- `whisper_print_timings` reports `encode time = ... ms / N runs`
  (`src/whisper.cpp:4288`), which is the honest place to observe the warm
  encoder cost — but it measures the whisper.cpp side of the call, and includes
  Core ML dispatch.

### 6.4. Compatibility-patch ownership

Every supported build binds three identities:

1. upstream whisper.cpp commit;
2. SHA-256 of the exact compatibility-patch bytes;
3. SHA-256 of the built `whisper-cli` binary.

The cache name includes the patch digest, the embedded bytes are hashed before
cache access, and the checkout is validated before and after the build. The
validator accepts exactly the declared diff on `src/whisper.cpp` and no other
tracked or untracked state. Direct tests cover hash drift, clean/unpatched and
modified trees, extra files, materialization, atomic publication, and staging
cleanup after a failed patch application.

Keep the patch until a proposed pin bump is inspected and physically tested.
Delete it only when the new upstream source initializes `whisper_batch` safely
on every early-return path, the unpatched candidate passes the repeated
negative control without a crash, ordinary GGML and Core ML positive paths
still pass, and manifests/cache identity are migrated deliberately. Current
upstream `2ca53bb…` still declares `whisper_batch batch;`, so the deletion
condition is not met.

Sanitizer builds are useful upstream-quality evidence, but do not overclaim
them: AddressSanitizer can catch the resulting invalid free, while detection of
the uninitialized read itself may require MemorySanitizer or platform-specific
instrumentation. The physical exact-exit/marker gate remains mandatory.

## 7. Parity measurement

### 7.1. Bind parity to the compiled artifact, at the runtime's compute units

The runtime loads a `.mlmodelc` with `MLComputeUnitsAll`
(`src/coreml/whisper-encoder.mm:29`, `:31`). Parity measured against an
`.mlpackage`, or at some other compute-unit setting, describes a configuration
that is not deployed. coremltools can load the exact compiled directory:

```python
import coremltools as ct

encoder = ct.models.CompiledMLModel(
    "/abs/path/to/ggml-model-encoder.mlmodelc",   # the published bytes
    compute_units=ct.ComputeUnit.ALL,             # matches whisper-encoder.mm:29
)
candidate = encoder.predict({"logmel_data": mel_fp32})["output"]
```

`CompiledMLModel` takes "the path to a compiled model directory, ending in
`.mlmodelc`" and accepts the same `ComputeUnit` enum as `MLModel`
(`coremltools/models/_compiled_model.py`). the implementationsitory currently predicts
through `ct.models.MLModel(<package>)` inside `_convert_and_predict()` in
`whisper_tuner/export/coreml.py`; switching the parity prediction to the
compiled path is an **evidence-binding** improvement (same bytes, same compute
units as the deployed run), not necessarily a numerical one.

The same applies to the export's `compute_units` option: exporting or measuring
under `cpu-only` produces a number that whisper.cpp will never reproduce,
because whisper.cpp cannot be told to use CPU only without patching
`whisper-encoder.mm`.

### 7.2. The reference side

Feed **the same** mel to both sides. Reference is the frozen Hugging Face
encoder, fp32, eval mode, on CPU, under `torch.inference_mode()`, with
`attn_implementation="eager"` so the traced graph and the reference agree.
Compare against `last_hidden_state` shaped `(1, 1500, hidden_size)`.

### 7.3. SNR is the right statistic, and Apple agrees

the reference implementation's frozen FP32 policy v2 is **SNR ≥ 100 dB AND cosine ≥ 0.9999**,
with max-absolute-error and mean-absolute-error recorded but non-gating. That
choice is methodologically consistent with coremltools' own converter test
harness, which validates numeric agreement with SNR/PSNR rather than an
absolute-error ceiling:

```python
def compute_snr_and_psnr(x, y):
    assert len(x) == len(y)
    eps = 1e-5
    eps2 = 1e-10
    noise = x - y
    noise_var = np.sum(noise**2) / len(noise)
    signal_energy = np.sum(y**2) / len(y)
    max_signal_energy = np.amax(y**2)
    snr = 10 * np.log10((signal_energy + eps) / (noise_var + eps2))
    psnr = 10 * np.log10((max_signal_energy + eps) / (noise_var + eps2))
    return snr, psnr
```

(`coremltools/converters/mil/testing_utils.py:867`–`877`.)

Two convention differences from the implementationsitory's
`signal_to_noise_ratio_db(reference, candidate)` are worth recording so the
numbers are comparable:

- coremltools treats the **second** argument as the signal; the implementationsitory
  treats the **first** (the PyTorch reference) as the signal. At ≥ 100 dB the
  two are numerically indistinguishable; at low SNR they are not.
- coremltools divides by element count and adds `eps`; the implementationsitory sums and
  handles the degenerate cases by clamping. The ratio is identical for
  same-length tensors.

Do not restate the policy version history here — it lives in
`whisper_tuner/resources/evaluation_suites.json`, including the recorded
failure of policy v1. The load-bearing rule for a new run is: the gate is
whatever the frozen resource says, and it is chosen **before** the measurement.

### 7.4. End-to-end WER methodology

Numeric parity on one probe is necessary, not sufficient. The end-to-end
comparison must:

1. run the **source Hugging Face model** on the frozen fixtures to produce the
   reference transcript;
2. run the **same pinned whisper.cpp binary** on the **same** `ggml-*.bin`
   *without* the `.mlmodelc` present, producing the GGML-encoder transcript;
3. run it again **with** the `.mlmodelc` present, producing the hybrid
   transcript, and assert the four load conditions of section 2.3;
4. compare normalized transcripts and report WER deltas for hybrid-vs-source
   and hybrid-vs-GGML.

Step 2 is what separates "Core ML changed the answer" from "conversion was
already wrong". It requires a **fallback-enabled** or separately built runtime,
or simply the non-Core-ML build of the same commit — do not enable
`WHISPER_COREML_ALLOW_FALLBACK` in the build used for step 3.

Persist hashes and bounded edit counts, not raw audio or raw transcripts, for
any fixture that is not already public.

## 8. Observing actual compute placement

The load marker says nothing about ANE, GPU, or CPU. Three independent
instruments, in decreasing order of evidentiary strength:

### 8.1. `MLComputePlan` (structural, per-operation)

```python
from coremltools.models.compute_plan import MLComputePlan
from coremltools import ComputeUnit

plan = MLComputePlan.load_from_path(
    "/abs/path/to/ggml-model-encoder.mlmodelc",
    compute_units=ComputeUnit.ALL,
)
# MLModelStructureProgram is not subscriptable on coremltools 9.0; functions
# are reached through its `functions` mapping (verified on this repo's pin).
program = plan.model_structure.program
for op in program.functions["main"].block.operations:
    usage = plan.get_compute_device_usage_for_mlprogram_operation(op)
    cost = plan.get_estimated_cost_for_mlprogram_operation(op)
    # usage.preferred_compute_device, usage.supported_compute_devices, cost.weight
```

- `load_from_path` takes an `.mlmodelc` directory — the same artifact
  whisper.cpp loads (`coremltools/models/compute_plan.py:403`–`448`).
- `MLComputePlanDeviceUsage` exposes `preferred_compute_device` and
  `supported_compute_devices` (`compute_plan.py:295`–`310`);
  `MLComputePlanCost.weight` is a `[0.0, 1.0]` share of total model execution
  (`compute_plan.py:313`–`323`).
- Device classes are `MLCPUComputeDevice`, `MLGPUComputeDevice`, and
  `MLNeuralEngineComputeDevice` (which also exposes `total_core_count`)
  (`coremltools/models/compute_device.py:68`–`135`).
- **Availability: macOS 14.4+ / iOS 17.4+** (Apple's `MLComputePlan` reference).
  `load_from_path` raises `ValueError("MLComputePlan is not supported.")` when
  the native proxy is unavailable.

What it proves: the framework's *plan*. It is a strong, cheap, scriptable
signal and it is per-operation, so it also identifies which ops break the ANE
chain. It is not a recording of an actual run.

### 8.2. Instruments Core ML template (behavioral)

The `Core ML` template is present in a stock Xcode install — locally verified
via `xcrun xctrace list templates` on macOS 15.7.7. It can be driven headlessly:

```bash
xcrun xctrace record \
  --template 'Core ML' \
  --output coreml-run.trace \
  --launch -- \
  ./build-coreml/bin/whisper-cli -m .../ggml-model.bin -f .../fixture.wav
```

What it proves: what actually executed, including per-model load and prediction
events. Cost: the `.trace` needs inspection, and automated extraction is more
work than `MLComputePlan`.

### 8.3. `powermetrics --samplers ane_power` (corroborating only)

Locally verified on macOS 15.7.7: `powermetrics` documents an `ane_power`
sampler — "dedicated rail ane power and frequency info" — and includes it in
both the `default` and `all` sampler groups. Nonzero ANE rail power during the
encoder window corroborates ANE execution.

Caveats that keep this out of the gate: it requires `sudo`, it is system-wide
rather than per-process, and Apple's own help text warns that "Average power
values reported by powermetrics are estimated and may be inaccurate".

### 8.4. The rule

Do not claim ANE residency from `ComputeUnit.ALL`, from a successful
transcription, from the load marker, or from a speedup. Claim it only from
`MLComputePlan` or an Instruments trace, and say which one.

## 9. Failure modes

| Symptom | Cause | Fix |
| --- | --- | --- |
| No Core ML lines at all; run succeeds | Binary built without `-DWHISPER_COREML=1`, or a stale non-Core-ML `whisper-cli` was invoked | Assert `COREML = 1`; build into a distinct directory; hash the binary |
| `failed to load Core ML model from '<path>'`, run continues | Build enabled `WHISPER_COREML_ALLOW_FALLBACK` | Rebuild with it `OFF`; assert the failure marker is absent |
| `failed to load ...`, exit `3` | `.mlmodelc` missing, misnamed, unreadable, or built for an incompatible OS | Recompute the derivation of section 4.1 against the literal `-m` string |
| Load marker present, transcript empty or nonsense | Prediction failed silently (`error:nil`) — usually a feature-name or shape mismatch | Validate input/output names and shapes on the Python side; compare against the GGML-encoder transcript |
| Load marker present, garbage only with `-ac N` | Mel width `2N` ≠ the fixed 3000-frame encoder input | Never pass `-ac` with a Core ML encoder |
| Gate log has no Core ML lines but exit `0` | `-np` / `--no-prints` silenced the log callback | Remove `-np` |
| N load markers for one file | `-p N` created N states, each loading Core ML | Use `-p 1` for gates |
| First positive run appears hung | ANE compiling the model to a device-specific format | Raise the first-run timeout; add a positive-path warm-up run |
| `xcrun coremlc` not found in CI | Only Command Line Tools installed | Install full Xcode; `xcrun --find coremlc` must resolve inside `Xcode.app` |
| `ane_transformers` install destroys the environment | Its `install_requires` pins `torch<=1.11.0` | Never install it into the project env; isolate it, or use the implementation's HF converter |
| Derived path truncated at a directory name | Model file has no extension and a parent directory contains a `.` | Always name the model `*.bin` |

## 10. What not to do

| Anti-pattern | Why it fails | Do this instead |
| --- | --- | --- |
| Grep stdout for the load marker | All logs go to stderr | Concatenate both streams, or capture stderr specifically |
| Use `-np` to keep gate output tidy | Silences the exact lines being asserted | Keep the log; hash it into evidence |
| Accept "nonzero exit" as the negative control | Path typos and bad args also exit nonzero | Require exit `3` plus both failure markers |
| Accept or retry `SIGABRT` | Converts undefined behavior into probabilistic success | Fail immediately; patch the uninitialized batch and rerun once from a known identity |
| Require a clean whisper.cpp checkout | The supported runtime deliberately carries one patch | Require the exact declared diff and no other change |
| Assert only `COREML = 1` | Proves the build, not the load | Assert both markers plus the absent-failure condition |
| Assert only `Core ML model loaded` | A fallback-enabled build could still have failed earlier in another run's log | Assert the full four-condition set from one run |
| Measure parity on the `.mlpackage` | Not the bytes the runtime loads | Use `CompiledMLModel` on the published `.mlmodelc` |
| Measure parity under `cpu-only` | whisper.cpp hardcodes `MLComputeUnitsAll` | Match the runtime's compute units |
| Gate on `np.allclose(..., atol=1e-2)` | Tolerance chosen by folklore, statistic chosen wrong | Use the frozen SNR + cosine policy |
| Claim ANE from a 3x speedup | Metal GPU also produces large speedups | `MLComputePlan` or Instruments, named explicitly |
| Rebuild the artifact after the gate passes | The published bytes were never tested | Gate the exact digest that ships |
| Reuse one build directory for both runtimes | Silent ordinary-encoder runs | Separate `build/` and `build-coreml/` |

## 11. Verified versus plausible

**Verified (primary source, cited by file and line above):** every log string
and its emission condition; the `ctx_coreml` → `whisper_encode_external` proof
chain; `WHISPER_COREML` / `WHISPER_COREML_ALLOW_FALLBACK` defaults and effects;
the `-encoder.mlmodelc` derivation algorithm; stderr as the log sink; `-np`
silencing; whisper-cli exit codes; the `[1, n_mels, 3000]` input and
`n_ctx * n_state` float output with its unchecked `memcpy`; hardcoded
`MLComputeUnitsAll`; `error:nil` on both Core ML calls; per-processor state
creation under `-p`; upstream converter's fixed 3000 shape and FP32 default;
`ane_transformers` pins and its 2022 last release; coremltools 9.0
`_TORCH_MAX_VERSION = 2.7.0` (warning only) and `_SKLEARN_MAX_VERSION = 1.5.1`;
`compute_snr_and_psnr`; `MLComputePlan` API surface and its macOS 14.4+
availability; `CompiledMLModel` accepting `.mlmodelc`; the Hugging Face
3000-frame check and its exact error text.

**Verified locally on macOS 15.7.7 (2026-07-30):** `coremlc` lives only inside
`Xcode.app`, not in Command Line Tools; Instruments ships a `Core ML` template;
`powermetrics` documents an `ane_power` sampler in its `default` and `all`
groups.

**Verified locally on 2026-08-01:** pinned and current upstream source leave
`whisper_batch` uninitialized across six pre-batch-init failure branches;
teardown unconditionally frees the member. The repeated unpatched/patch A/B in
section 6.2 produced 15 crashes in 60 unpatched attempts and zero crashes in
200 patched attempts. Dedicated ordinary-GGML and Core ML hybrid gates passed
with upstream, patch, and binary identities recorded. The exact physical nodes
must replay after the final checkout is frozen.

**Plausible (community reports, not primary):** first-load ANE compile times in
the tens of seconds for larger encoders; `ANECompilerService` pegging a core or
appearing to hang during that compile; the `kill -9 ANECompilerService`
workaround in the historical guide. These are consistent with upstream's own
"first run is slow" statement, but no primary source bounds the latency. Treat
any specific number as an observation from one machine.

**Not established, and deliberately not claimed here:** that the implementationsitory's
recorded FP32 residual is caused by the torch-version warning, by compute-unit
placement, or by any specific coremltools defect. The debug notes record that
Torch 2.7, every compute-unit mode, and multiple deployment targets all failed
the same way.

## Works Cited

1. [whisper.cpp `src/whisper.cpp` @ `97c56f1`](https://github.com/ggml-org/whisper.cpp/blob/97c56f1dc1d1100a9d859c865a20c82d22f823ed/src/whisper.cpp) — Core ML load block, log strings, path derivation, external-encoder predicate, mel tensor, timings, default log sink. Retrieved 2026-07-30.
2. [whisper.cpp `src/whisper.cpp` @ current checked `master` (`2ca53bb`)](https://github.com/ggml-org/whisper.cpp/blob/2ca53bb45e38748d07b310eeb36245a7157ac882/src/whisper.cpp) — the same uninitialized-batch lifecycle remains as of 2026-08-01.
3. [whisper.cpp `src/coreml/whisper-encoder.mm`](https://github.com/ggml-org/whisper.cpp/blob/4523d0ce373ee4b2176b3251fff29fd4864fcf38/src/coreml/whisper-encoder.mm) — `MLComputeUnitsAll`, `error:nil`, `MLMultiArray` shape/strides, output `memcpy`. Retrieved 2026-07-30.
4. [whisper.cpp `src/coreml/whisper-encoder-impl.m` / `.h`](https://github.com/ggml-org/whisper.cpp/blob/4523d0ce373ee4b2176b3251fff29fd4864fcf38/src/coreml/whisper-encoder-impl.h) — `logmel_data` / `output` feature names, documented `1 × 80 × 3000` input. Retrieved 2026-07-30.
5. [whisper.cpp `src/CMakeLists.txt`](https://github.com/ggml-org/whisper.cpp/blob/4523d0ce373ee4b2176b3251fff29fd4864fcf38/src/CMakeLists.txt) — `WHISPER_USE_COREML` definition, framework `FATAL_ERROR`, `whisper.coreml` target. Retrieved 2026-07-30.
6. [whisper.cpp `CMakeLists.txt`](https://github.com/ggml-org/whisper.cpp/blob/4523d0ce373ee4b2176b3251fff29fd4864fcf38/CMakeLists.txt) — `WHISPER_COREML` and `WHISPER_COREML_ALLOW_FALLBACK` default to `OFF`. Retrieved 2026-07-30.
7. [whisper.cpp `examples/cli/cli.cpp`](https://github.com/ggml-org/whisper.cpp/blob/4523d0ce373ee4b2176b3251fff29fd4864fcf38/examples/cli/cli.cpp) — exit codes, `system_info` print, `--no-prints` log disable, `--audio-ctx`. Retrieved 2026-07-30.
8. [whisper.cpp `models/generate-coreml-model.sh`](https://github.com/ggml-org/whisper.cpp/blob/4523d0ce373ee4b2176b3251fff29fd4864fcf38/models/generate-coreml-model.sh) — conversion, `xcrun coremlc compile`, rename to `ggml-<name>-encoder.mlmodelc`. Retrieved 2026-07-30.
9. [whisper.cpp `models/convert-whisper-to-coreml.py`](https://github.com/ggml-org/whisper.cpp/blob/4523d0ce373ee4b2176b3251fff29fd4864fcf38/models/convert-whisper-to-coreml.py) — fixed `(1, n_mels, 3000)` input, `mlprogram`, FP32 default, ANE reshaping, SDPA disable. Retrieved 2026-07-30.
10. [whisper.cpp README, "Core ML support"](https://github.com/ggml-org/whisper.cpp/blob/4523d0ce373ee4b2176b3251fff29fd4864fcf38/README.md#core-ml-support) — documented build, expected log lines, first-run compile statement, Xcode and Python 3.11 guidance. Retrieved 2026-07-30.
11. [whisper.cpp issue #2278, "When built with CoreML, can no longer run normal models"](https://github.com/ggml-org/whisper.cpp/issues/2278) — community-visible consequence of fail-closed init; open since 2024-07-03. Retrieved 2026-07-30.
12. [coremltools `coremltools/_deps/__init__.py` @ 9.0](https://github.com/apple/coremltools/blob/9.0/coremltools/_deps/__init__.py) — `_TORCH_MAX_VERSION`, `_SKLEARN_MAX_VERSION`, warning-only enforcement. Retrieved 2026-07-30.
13. [coremltools `coremltools/converters/mil/testing_utils.py` @ 9.0](https://github.com/apple/coremltools/blob/9.0/coremltools/converters/mil/testing_utils.py) — `compute_snr_and_psnr`. Retrieved 2026-07-30.
14. [coremltools `coremltools/models/compute_plan.py` @ 9.0](https://github.com/apple/coremltools/blob/9.0/coremltools/models/compute_plan.py) — `MLComputePlan.load_from_path`, `MLComputePlanDeviceUsage`, `MLComputePlanCost`. Retrieved 2026-07-30.
15. [coremltools `coremltools/models/compute_device.py` @ 9.0](https://github.com/apple/coremltools/blob/9.0/coremltools/models/compute_device.py) — CPU/GPU/Neural Engine device classes. Retrieved 2026-07-30.
16. [coremltools `coremltools/models/_compiled_model.py` @ 9.0](https://github.com/apple/coremltools/blob/9.0/coremltools/models/_compiled_model.py) — `CompiledMLModel` loads an `.mlmodelc` with a chosen `ComputeUnit`. Retrieved 2026-07-30.
17. [coremltools 9.0 release notes](https://github.com/apple/coremltools/releases/tag/9.0) — published 2025-11-10; "Support for PyTorch 2.7"; Python 3.13 support; `macOS26` targets. Retrieved 2026-07-30.
18. [Apple Developer, `MLComputePlan`](https://developer.apple.com/documentation/coreml/mlcomputeplan) — availability macOS 14.4+, iOS 17.4+. Retrieved 2026-07-30.
19. [Hugging Face Transformers, `modeling_whisper.py`](https://github.com/huggingface/transformers/blob/main/src/transformers/models/whisper/modeling_whisper.py) — `expected_seq_length = max_source_positions * conv1.stride * conv2.stride` and its exact `ValueError`. Retrieved 2026-07-30.
20. [`apple/ml-ane-transformers` `setup.py`](https://github.com/apple/ml-ane-transformers/blob/main/setup.py) — `torch>=1.10.0,<=1.11.0`, `protobuf<=3.20.1`. Retrieved 2026-07-30.
21. [`ane_transformers` on PyPI](https://pypi.org/project/ane-transformers/) — 0.1.3 uploaded 2022-08-09; sdist only. Retrieved 2026-07-30.
22. [whisper.cpp issue #2750](https://github.com/ggml-org/whisper.cpp/issues/2750) — independent Core ML load-failure and invalid-free symptom report; it does not identify the root cause. Retrieved 2026-08-01.
</content>
</invoke>
