MissingManual

Whisper / ANE graph surgery · Collective Library

Core ML Whisper Encoder Compression and ANE Graph Surgery

Collective Library edition. This is the complete technical report. Private filesystem paths, internal run identifiers, campaign-control notes, and repository navigation were removed. Technical claims, code, measurements, evidence labels, citations, corrections, and falsification criteria are preserved.

How to test low-bit weight storage and graph rewrites on a fixed-shape Whisper encoder without confusing a smaller package, a compiler preference, or a fast toy with a faster accurate model on an M1.

Provenance. Ingested 2026-08-08 from a Gemini Deep Research Max run:

  core-ml-whisper-encoder-compression-and-ane-graph-surgery/
  2026-08-09T06-15-24-438Z/raw-report.md

SHA-256 77d6b4276a03989cba8920f9e06ed356c892b1d0a900bc5484c8371456700a7a

The Core ML Tools 9 API surface was re-checked against Apple's documentation via Context7. Sections marked Implementation evidence are firsthand measurements and take precedence over the external draft where they conflict.

The boundary that matters

The deployed large-v3-turbo hybrid has three materially different costs:

fixed 3000-frame log-mel input
        |
        v
Core ML encoder  ---- [1, 1500, 1280] ----> whisper.cpp decoder
        |                                      |
  weight traffic,                        autoregressive token
  attention, placement                   loop and KV cache

Compressing the Core ML package changes encoder weight storage. It does not remove decoder tokens, audio preprocessing, VAD, or process startup. The current campaign therefore requires an encoder below roughly 237 ms if the measured CPU decoder and remaining compute stay around 263 ms. A package-size win is useful only if the same compiled graph crosses that mechanical latency threshold.

Implementation evidence. The fastest sealed M1 baseline is a 622.045 ms encoder and 884.784 ms warmed compute total on a Macmini9,1. The encoder alone is already 122.045 ms over the entire 500 ms goal. Residency and launch-overhead work cannot close that gap.

What Core ML Tools actually provides

coremltools.optimize.coreml.palettize_weights rewrites constant weights in an mlprogram into lookup-table storage. It is post-training and data-free: it does not calibrate activations or establish ASR quality.

import coremltools as ct
import coremltools.optimize as cto

source = ct.models.MLModel("encoder.mlpackage", skip_model_load=True)
op_config = cto.OpPalettizerConfig(
    mode="uniform",
    nbits=4,
    granularity="per_grouped_channel",
    group_size=16,
    weight_threshold=2048,
)
config = cto.OptimizationConfig(
    global_config=None,
    op_type_configs={"linear": op_config, "conv": op_config},
)
candidate = cto.palettize_weights(source, config)
candidate.save("encoder-lut4.mlpackage")

Core ML Tools 9 documents these relevant choices:

Control Documented behavior Experimental consequence
mode kmeans, uniform, unique, or custom Mode is part of artifact identity; never change it inside an A/B.
nbits 1, 2, 3, 4, 6, or 8 for supported modes Four bits means a 16-entry palette, not integer matrix math.
granularity per_tensor or per_grouped_channel More palettes can improve fidelity but may reduce runtime performance.
group_size Used only for grouped-channel palettes Group 16 is a hypothesis imported from a different graph, not a Whisper default.
weight_threshold Defaults to 2048 elements Small constants remain uncompressed; record coverage and bytes.
cluster_dim Values above one enable vector palettes Treat as a separate, newer-runtime experiment.

Grouped-channel and vector features require the iOS 18/macOS 15 operation set. A source mlprogram converted for macOS 14 cannot be assumed to accept them. Export the experimental baseline and candidate from the same source model with the same macOS 15 target before comparing them.

Palettization stores weights compactly and expands them through constexpr_lut_to_dense. It does not prove that prediction executes four-bit math, that the Neural Engine reads four times fewer bytes, or that latency falls with package size. Those are physical questions.

The op-type filter is deliberate. Core ML Tools follows each constant to its consumer when choosing a compression config. Restricting the candidate to linear and conv leaves the positional embedding and other add-path constants alone. A global config would risk quantizing the 1500x1280 positional table for little size benefit and a disproportionate timestamp/quality risk.

The useful Plateau transfer—and its limit

Implementation evidence from the sibling Plateau repository. On its fixed-shape diffusion-transformer graph, uniform LUT4 with grouped-channel palettes and group size 16 compressed a four-block weight blob from about 1,745 MB to 436 MB while preserving Neural Engine preference for every material linear. A different representation—blocked INT4 from linear_quantize_weights, block size 32—moved all 28 linears to CPU even though the top-line report still said 81.9% Neural Engine. On an M1 Mini, the LUT4 arm reported 99.74% preferred Neural Engine placement.

The transferable lesson is about representation, not the word “grouped”:

  • grouped-channel palette compression was the positive arm;
  • grouped/blockwise linear quantization was the negative arm;
  • top-line device percentages hid the material CPU fallback;
  • the result used another architecture, tensor geometry, and in some legs random weights, so it says nothing yet about Whisper quality or latency.

Plateau also found only 1.03x aggregate throughput from two concurrent ANE processes. Parallel encoders are therefore a poor primary latency bet, although that result remains topology-specific.

Claims the raw research did not establish

The external report contained several strong prescriptions that are not safe to promote as facts:

  • “Every linear must become a 1x1 convolution.” False as a universal rule. Layout-sensitive rewrites can help some graphs, but the existing Whisper package already runs and produces strong ANE power evidence. Inspect the exact compiled plan before rewriting trained math.
  • “A two-gigabyte cliff requires chunking this encoder.” Plateau measured a cliff on another graph. The deployed Whisper encoder is already below that boundary. Chunking adds dispatches and intermediate transfers and needs its own A/B.
  • “M1 lacks blocked-quantization kernels while M3/M4 accelerate them.” The public Core ML API documents operation-set availability, not this per-chip scheduling guarantee. Keep the Plateau result device/build-scoped.
  • “The ANE has a four-megabyte shared L2 cache.” The supplied sources did not establish a public architectural contract. Do not derive Whisper chunk sizes from it.
  • “Change LayerNorm epsilon, masks, or softplus, then calibrate.” Do not change trained semantics without a reproduced numerical failure on the exact Whisper graph. None of those edits is required for weight-only palettization.
  • “Quantization cannot reach the target.” Also unproven. The current encoder needs a 2.62x latency improvement to fit the residual budget. A fourfold storage reduction is at least large enough to falsify before paying for architectural distillation.

Local graph inventory. The current compiled Whisper encoder contains 192 linear, 64 matmul, 32 softmax, and two conv operations, with no fused SDPA operation. Its weight blob is 1,273,971,776 bytes—well below Plateau's separate 2^31-byte cliff. The first LUT4 probe should therefore expect exactly 194 LUT operations for the linear/conv weights, zero LUTs for the positional embedding, and no attention or chunking rewrite.

Cheapest-falsifier ladder

Run the ladder in order. A failure stops the arm; it does not authorize a more complex rewrite.

1. Package physics

Build one same-source macOS 15 pair:

  • FP16 baseline;
  • LUT4 uniform, grouped-channel, group 16, threshold 2048.

Record source tree digest, Core ML Tools version, full config, package and compiled-tree digests, bytes, operation histogram, number of constexpr_lut_to_dense operations, and palette cardinality. Kill if conversion or compilation fails, the ABI changes, the candidate does not contain exactly 194 linear/conv LUT operations, any positional-embedding constant is palettized, or the selected weight bytes are not at least 3.5x smaller.

Do not compile this candidate through the production compile_package helper. That helper correctly binds only Float16 and Float32 storage and rejects an unknown compiled storagePrecision. Use raw xcrun coremlc compile, label the receipt experimental_non_certified, archive the compiled metadata verbatim, and record storage format (lut4) separately from compute precision (fp16).

2. Static placement

Load MLComputePlan from the compiled .mlmodelc under CPU_AND_NE. Apple's API reports anticipated device usage and estimated cost per operation. Inspect material linear, matmul, convolution, and attention operations individually; do not accept an aggregate percentage.

Kill if any material candidate operation moves from a baseline Neural Engine preference to CPU/GPU. A static plan is compiler intent, not runtime proof.

3. Compiled numerical parity

Run the exact compiled bundles through the same three frozen encoder probes:

  • real speech fixture;
  • processor-produced silence/padding;
  • deterministic stress tensor.

Record max and mean absolute error, cosine similarity, and SNR per probe. Do not borrow the production FP16 gate for LUT4 by relabeling precision. The quantized arm needs a new experimental policy plus end-to-end transcript gates.

4. End-to-end quality

Use the same decoder bytes and greedy settings for both encoders. Require:

  • no text regression on the public smoke and frozen pathology fixtures;
  • no WER regression beyond a preregistered bound on held-out speech;
  • timestamp and no-speech checks, because those can fail before normalized text;
  • forbidden-log checks proving Core ML did not silently fall back.

5. Physical M1 latency

Run warmed, paired, rotating A/B trials on a quiet M1. Pair MLComputePlan with runtime rails or an Instruments trace. Measure encoder, decoder, preprocessing, and total separately for both approximately 15-second and 30-second clips.

The campaign go gate is mechanical:

candidate encoder median < 237 ms
candidate warmed compute p50 < 500 ms
candidate warmed compute p95 < 550 ms
quality gates all pass

If LUT4 preserves quality and placement but misses the latency gate, its bytes may still be useful in combination with a trained smaller encoder. It does not win this campaign by itself.

What to try only after LUT4

Candidate Why it might work Primary risk First kill check
Per-tensor LUT4 Fewer palettes and simpler expansion Worse weight fidelity Frozen-probe and transcript parity
Group-16 kmeans LUT4 Better fit than uniform Slow conversion; no runtime win Same placement and bytes, then parity
INT8 per-channel Conservative compression Only about 2x bytes, likely insufficient alone Encoder must beat FP16 materially
Mixed FP16/LUT4 by op Protects outlier-sensitive weights Branches or exclusions erase bandwidth win Package coverage and plan first
Explicit attention rewrite Can bypass a reproduced fused-op defect Semantic drift and graph explosion One-layer numerical probe
Trained encoder compression Can remove compute, not just bytes Data, training, and generalization cost Teacher-student null arm

Do not build an automatic search over this matrix first. One exact artifact, one exact M1, and one hard kill gate provide more information than a fleet of unqualified package-size measurements.

References

  1. Core ML Tools palettization API
  2. Core ML Tools grouped-channel palettization example
  3. Core ML Tools ML Program utilities
  4. Core ML Tools MLComputePlan
  5. Distil-Whisper paper