# Core ML Fused SDPA for Whisper-Shaped Encoders on Base M1

> **Collective Library edition.** This is the complete technical report. Private filesystem paths, internal run identifiers, campaign-control notes, and repository navigation were removed. Technical claims, code, measurements, evidence labels, citations, corrections, and falsification criteria are preserved.

How to construct and falsify a fused scaled-dot-product-attention Core ML
experiment without mistaking a fused MIL spelling, requested compute units, or
a residual-path output for correct Neural Engine execution.

**Provenance.** Ingested 2026-08-09 from a Gemini Deep Research Max run:

```text
  core-ml-fused-scaled-dot-product-attention-for-whisper-shaped-encoders-on-base-a/
  2026-08-09T09-09-20-079Z/raw-report.md

SHA-256 df7173b5c5500aa980a09af79b703b6fa9929bd78104743a6cf1dde6e463b3d5
```

The external synthesis was corrected during ingest against the installed Core
ML Tools 9 source, Apple's MIL operation reference, PyTorch's SDPA contract,
and firsthand Plateau device receipts. Sections marked **Implementation evidence** take
precedence where they conflict with the raw report.

## Executive Summary

Core ML Tools 9 can preserve PyTorch
`torch.nn.functional.scaled_dot_product_attention` as the iOS 18 MIL
`scaled_dot_product_attention` operation when all of these conditions hold:

- the conversion target is `ct.target.macOS15` or the corresponding iOS 18
  operation set;
- Q, K, and V have equal rank and rank at least three after frontend handling;
- dropout is zero;
- PyTorch's `scale` argument is omitted or `None`.

The last condition is the easy one to miss. The current Transformers Whisper
implementation multiplies Q by `head_dim**-0.5`, then calls its attention
backend with `scaling=1.0`. Core ML Tools 9 accepts that explicit scale, but
deliberately decomposes the operation into matmul, softmax, and matmul. The
fusion candidate must instead remove the Q pre-scale and call SDPA with raw Q,
K, and V while omitting `scale`. For head dimension 64, PyTorch and MIL then
apply the default `1/sqrt(64) = 0.125` factor. The math is algebraically equal;
the FP16 operation order is not, so paired numerical and ASR parity remain
mandatory.

There is no public or local law requiring 64 attention heads for Neural Engine
placement. **Implementation evidence:** the sibling Plateau project placed fused SDPA on
a base M1 under macOS 15.7.7 at S=288, 32 heads, and head dimension 128 after
three linear projections. Its first, more minimal graph exposed Q/K/V directly
as model inputs and placed SDPA on CPU. The same fused-op family returned all
zeros when Neural Engine-placed on macOS 26.5.2. Placement is therefore graph-
and OS-build-specific, and it does not prove correctness.

Plateau has no attention-only latency result. It establishes an existence
proof and two severe failure modes, not a speedup. The useful next measurement
is an attention-output A/B at the exact d750 internal shape, `[2, 20, 750, 64]`,
with identical trained projections. Expose the attention output directly;
testing only a residual block can hide a dead zero-valued attention path.

## Evidence Boundary

| Claim | Status | Consequence |
| --- | --- | --- |
| Fused MIL SDPA exists in the iOS 18/macOS 15 operation set | Documented | Target macOS 15 for both A/B arms. |
| Explicit `scale=1.0` converts in Core ML Tools 9 | Locally source- and conversion-verified | It decomposes; successful conversion is not fusion. |
| Boolean and float masks are accepted by the MIL op | Documented | Preserve PyTorch mask polarity and broadcasting exactly. |
| Fused SDPA can place on a base-M1 Neural Engine | Firsthand at S=288, H=32, Dh=128 | Exact H=20, S=750 placement remains unknown. |
| Fused SDPA is faster than explicit attention on M1 | No local evidence | Require paired warmed latency; do not project a percentage. |
| ANE-placed fused SDPA is correct on macOS 26 | Refuted on one M2 Ultra, macOS 26.5.2 | OS build is part of artifact qualification. |
| A 64-head threshold controls placement | Refuted as a universal rule | Do not spoof heads or duplicate tensors around folklore. |

The external report also cited MPS 32-bit indexing failures at enormous tensor
sizes. That is a different backend and scale regime; it is not evidence about
Core ML at S=750. Likewise, reported FP16 exponential cliffs and private
ANEForge heuristics are practitioner clues, not public hardware contracts.

## The Exact Fusion Contract

### What Core ML Tools 9 does

Installed source in
`coremltools/converters/mil/frontend/torch/ops.py` makes the decision directly:

```python
can_use_fused_sdpa = (
    is_current_opset_version_compatible_with(target.iOS18)
    and scale is None
)
```

When that predicate is true, the frontend emits
`mb.scaled_dot_product_attention`. Otherwise it calls
`_decompose_scaled_dot_product_attention`. This has several consequences:

| Source spelling | macOS 15 result in Core ML Tools 9 |
| --- | --- |
| `F.scaled_dot_product_attention(q, k, v)` | Eligible for one fused MIL SDPA op. |
| Same call with `scale=None` | Eligible for one fused MIL SDPA op. |
| Same call with `scale=1.0` | Valid conversion, but decomposed. |
| Manual `softmax(q @ k.T * scale) @ v` | Explicit primitives by construction. |
| `dropout_p != 0` | Conversion rejected. |
| `attn_mask` plus `is_causal=True` | Frontend rejects the combination. |
| Pre-macOS-15 operation set | Decomposed because fused MIL SDPA is unavailable. |

The iOS 18 MIL operation itself has no `scale`, `dropout_p`, or `is_causal`
input. The PyTorch frontend can construct a causal boolean mask before calling
the fused operation, but the encoder candidate needs neither causality nor a
mask. Do not add mask machinery to the first d750 falsifier.

### Correct Whisper rewrite

The existing Transformers ordering is conceptually:

```python
q = q_proj(x) * (head_dim**-0.5)
out = F.scaled_dot_product_attention(
    q,
    k_proj(x),
    v_proj(x),
    dropout_p=0.0,
    is_causal=False,
    scale=1.0,
)
```

That preserves the historical Whisper rounding order but forces Core ML Tools
9 to decompose. The fusion candidate is:

```python
q = q_proj(x)
out = F.scaled_dot_product_attention(
    q,
    k_proj(x),
    v_proj(x),
    dropout_p=0.0,
    is_causal=False,
)
```

Do not keep the Q pre-scale while omitting `scale`; that would apply 0.125
twice. Do not pass `scale=1.0` and claim the resulting graph is fused. Inspect
the MIL program after conversion and require exactly one
`scaled_dot_product_attention` in the candidate and zero in the explicit arm.

Pin Torch and Core ML Tools in the receipt. The local Core ML Tools 9 install
warns that its most recently tested Torch is 2.7 while this checkout uses Torch
2.10; a passing conversion on this exact environment is evidence, but not a
portable compatibility guarantee.

## Shapes and Masks

Apple's MIL definition accepts Q, K, and V with rank at least three. Their ranks
must match; their batch dimensions must be broadcast-compatible; Q and K must
share the embedding dimension; and K and V must share source length. The raw
report's statement that tensors must be exactly `(1, 20, 750, 64)` was too
narrow.

The real d750 encoder folds two independent 750-position halves into the batch
axis. Its attention tensors are therefore:

```text
Q, K, V: [2, 20, 750, 64]
output:  [2, 20, 750, 64]
mask:    none
```

For graphs that do use a mask:

- boolean `True` means the element participates in attention;
- floating masks are added to the score;
- the mask must broadcast to the Q/K score shape;
- boolean and floating masks are not interchangeable without preserving
  polarity and sentinel behavior;
- do not pair an explicit attention mask with `is_causal=True` in the
  Core ML Tools PyTorch frontend.

The external report's blanket recommendation to prefer boolean masks was not
established. Use the source model's exact semantics, then test a mask-off/mask-
on pair. Avoid `NaN` sentinels. Treat `-inf` versus a large finite negative as a
separate numerical experiment rather than silently changing it.

## Plateau Firsthand Corrections

The public synthesis made strong scheduling claims that the local physical
receipts directly correct.

### Base M1, macOS 15.7.7

The sibling Plateau probe used random weights and this graph boundary:

```text
[1, 288, 4096]
  -> Q/K/V linear projections
  -> reshape + transpose to [1, 32, 288, 128]
  -> fused SDPA
  -> exposed attention output
```

On a `Macmini9,1` base M1, macOS 15.7.7 build `24G720`, Core ML Tools 9:

- the fused additive-mask SDPA operation was assigned to the Neural Engine;
- CPU-and-Neural-Engine output cosine versus Torch was approximately
  `0.999999405`, with max absolute error approximately `0.000504`;
- changing the mask produced a nonzero output change;
- explicit attention also placed its material math on the Neural Engine and
  produced correct output.

This is enough to refute a universal “at least 64 heads” requirement because
the measured graph had 32 heads. It is not a receipt for 20 heads or S=750.

The first Plateau harness passed Q/K/V as three model inputs. Core ML placed
the fused operation on CPU, and its correct outputs initially looked like a
successful ANE test. Adding the three projections changed placement. The
minimal faithful probe must therefore retain the producer context that induces
the desired schedule.

Firsthand source paths in the sibling repository:

```text
plateau/toys/coreml_sdpa_mask_probe/build_probe.py
plateau/toys/coreml_sdpa_mask_probe/run_probe.py
plateau/toys/coreml_sdpa_mask_probe/results_mini_m1.json
plateau/README/Notes/plateau-e4-ane-placement-probe-notes-2026-08-06.md
```

### macOS 26.5.2 zero-output failure

On one M2 Ultra running macOS 26.5.2, an ANE-placed fused SDPA returned all
zeros at S=64 and S=288, including the no-mask case. Boolean and causal forms
also failed. An additive-mask variant appeared correct only when Core ML did
not place SDPA on the Neural Engine. Explicit attention remained correct and
ANE-placed.

This failure explains why the falsifier must expose attention directly. A
residual block can return plausible nonzero output even if its attention branch
is exactly zero. It also means OS qualification is not optional: the same
package family can be correct on macOS 15 and wrong on macOS 26.

Plateau's SDPA correctness probe records no warm latency samples. Do not cite
its placement or energy campaigns as evidence that fused attention beats
explicit attention.

## Placement and Runtime Evidence

Requested `CPU_AND_NE` is a scheduler constraint, not proof that the operation
ran on the Neural Engine. A promoted result needs all of these layers:

1. **Source identity:** exact weights, wrapper source, Torch/Core ML Tools
   versions, conversion target, and source hash.
2. **MIL identity:** one fused SDPA in the candidate; explicit matmul-softmax-
   matmul in the control.
3. **Compiled identity:** package and `.mlmodelc` tree digests plus OS build.
4. **Static placement:** per-operation `MLComputePlan` device usage and
   estimated cost under the same compute-unit policy.
5. **Runtime corroboration:** synchronized signposted predictions with the Core
   ML Instruments template or supported per-rail telemetry.
6. **Correctness:** direct attention output compared against the same Torch
   reference and the paired explicit Core ML arm.

`MLComputePlan` describes anticipated device usage. It does not prove runtime
execution or numerical correctness. Keep the loaded `MLModel` alive while
loading a compute plan from its compiled path; Plateau found that releasing the
model deletes the temporary `.mlmodelc`. Loading a plan directly from an
`.mlpackage` is also the wrong boundary.

Do not claim that Metal System Trace alone proves ANE execution. Use the Core
ML instrument or ANE-specific evidence where the OS exposes it, and include a
CPU-only negative control. For `powermetrics`, enumerate supported samplers on
the target host instead of assuming an `ane_power` sampler exists.

## Smallest Physical Falsifier

Compare attention subgraphs, not full residual blocks. Both arms must target
macOS 15; changing the deployment target inside the A/B confounds the result.

### Frozen inputs and weights

- Input `x`: FP16 `[2, 750, 1280]` from the exact d750 layer boundary.
- Projections: identical trained Q, K, V weights and biases.
- Head view: `[2, 20, 750, 64]`.
- Output: expose `[2, 20, 750, 64]` before output projection and residual.
- Fixtures: deterministic stress tensor, processor-produced silence, and one
  frozen real-speech activation.

### Arms

| Arm | Attention spelling | Required MIL identity |
| --- | --- | --- |
| Explicit | `softmax((q * 0.125) @ k.T) @ v` in the current order | Matmul, softmax, matmul; no fused SDPA |
| Fused | raw Q/K/V into `F.scaled_dot_product_attention`, scale omitted | Exactly one fused SDPA |
| CPU-only negative | Same fused package under CPU-only | Correctness discriminator, not latency contender |
| Dead-path negative | Zero the attention output before export | Parity gate must reject it decisively |

Use the same macOS 15 target, precision, inputs, weights, I/O contract, host,
runner, warmup, and randomized paired trial order. Do not add a mask, residual,
MLState, palettization, or output projection until the primitive answers the
placement, correctness, and latency questions.

### Gates

Run the cheapest checks first:

1. **Conversion:** fused arm contains exactly one fused MIL operation. Kill if
   it decomposes, conversion fails, or the ABI differs.
2. **Static placement:** fused SDPA and producer projections prefer the Neural
   Engine. Kill the M1 ANE hypothesis if SDPA prefers CPU/GPU or is unknown
   while the explicit material ops prefer the Neural Engine.
3. **Direct parity:** all outputs finite; cosine at least `0.9999`; max-absolute
   error no worse than a preregistered multiple of the explicit Core ML arm;
   dead-path negative rejected. Any all-zero output is an immediate kill.
4. **Sensitivity:** a deterministic single-token perturbation produces a
   nonzero, reference-consistent response. If a mask is later added, its on/off
   pair must also change output in the expected direction.
5. **Warm latency:** paired, rotating A/B predictions on a quiet base M1. Keep
   p50, p95, raw samples, and compilation/load time separate.

There is no public latency prior strong enough to predeclare a universal 1.5x
speedup. For the active campaign, continue only if the measured per-layer delta
multiplied by the exact number of rewritten layers can close the remaining
end-to-end gap with margin. If the confidence interval includes zero or the
projected aggregate prize is below 5 ms, stop before building a full encoder.

## Failure Decision Tree

```text
No fused MIL op?
  -> Confirm macOS15 target and scale=None.
  -> If scale=1.0 is present, decomposition is expected.

Fused MIL op but SDPA not assigned to ANE?
  -> Confirm the three projections remain inside the graph.
  -> Compare the exact explicit arm under the same target and compute units.
  -> Do not spoof head count before the exact H20/S750 result exists.

ANE placement but zero or constant output?
  -> Stop immediately.
  -> Re-run CPU-only and explicit-attention controls.
  -> Record OS build; do not hide the failure behind a residual.

Correct output but no speedup?
  -> Accept the result: fusion is not a useful M1 lever on this graph/build.
  -> Do not add masks, slicing, private APIs, or a full-model rewrite.

Speedup but parity drift?
  -> Attribute first to changed FP16 scaling/fusion order.
  -> Inspect layer-local error and end-to-end transcript/timestamp behavior.
  -> Promote only under a separate topology-change quality policy.
```

## Later Experiments, Only After the Exact Probe

Core ML Tools exposes a `scaled_dot_product_attention_sliced_q` graph pass for
large query lengths. It replaces fused SDPA with a sliced implementation and
can reduce peak memory, but it is a different graph and may trade dispatches
for memory. Test it only if the unsliced S=750 fused arm is correct, placed, and
memory-limited.

Private ANEForge-style execution can probe hardware behavior, but undocumented
frameworks and daemon interfaces are not a production dependency. MLX fused
attention is relevant to a GPU architecture; it does not answer whether the
fixed-shape Core ML encoder improves on the M1 Neural Engine. Neither belongs
in the first experiment.

Do not preemptively rewrite linears as convolutions, split or duplicate heads,
change LayerNorm epsilon, add subtract-max softmax code, or chunk the entire
encoder. Each adds a variable before the primitive has produced any latency
evidence.

## Sources

1. [Core ML Tools iOS 18 MIL SDPA reference](https://apple.github.io/coremltools/source/coremltools.converters.mil.mil.ops.defs.html#coremltools.converters.mil.mil.ops.defs.iOS18.transformers.scaled_dot_product_attention)
2. [Core ML Tools MIL transformer graph passes](https://apple.github.io/coremltools/source/coremltools.converters.mil.mil.passes.defs.html#transformer)
3. [PyTorch functional SDPA](https://docs.pytorch.org/docs/stable/generated/torch.nn.functional.scaled_dot_product_attention)
4. [Apple `MLComputePlan`](https://developer.apple.com/documentation/coreml/mlcomputeplan-1w21n)
5. [Core ML Tools source](https://github.com/apple/coremltools)
