External-Encoder, Decoder-Only Whisper Runtime on Apple Silicon
Collective Library edition. This is the complete technical report. Private filesystem paths, internal run identifiers, campaign-control notes, and repository navigation were removed. Technical claims, code, measurements, evidence labels, citations, corrections, and falsification criteria are preserved.
August 12, 2026
Executive Summary
Upstream whisper.cpp documents a hybrid configuration in which Core ML runs the encoder while the decoder remains in the GGML model. It does not document a decoder-only GGML container for this pinned loader. Structural encoder deletion is a local experimental derivative and must be qualified as such.
The opportunity is nevertheless concrete. The ordinary file initializer in
the pinned runtime opens the model with std::ifstream. The loader declares
the expected tensors, allocates backend buffers for them, then reads tensor
bytes into those allocations. It is not an mmap-backed model path. Enabling
Core ML changes encoder compute, but by itself does not delete the native
encoder tensors from the model allocation.
The smallest safe decoder-only design therefore has four parts:
- derive an authenticated artifact that preserves the original header, mel filters, vocabulary, and every decoder record byte-for-byte while omitting every native encoder record;
- make the runtime declare and expect only those retained decoder tensors;
- make state creation and encoder execution fail closed unless the exact external Core ML sidecar loads and returns the required tensor; and
- prove exact tokens, text, EOT behavior, memory deletion, and repeated-call lifecycle behavior against the complete-model control.
Do not mix this lossless structural experiment with further quantization, cache shrinking, logits-API changes, or Core ML compression. Each changes a different causal mechanism.
1. Corrected Evidence Boundary
| Raw-report claim | Corrected verdict |
|---|---|
The pinned loader relies on a model mmap |
False. The ordinary file path uses std::ifstream; the model loader allocates backend tensors and reads bytes into them. |
| Deleting encoder file bytes automatically saves the same resident bytes | Not yet proven. It prevents those tensors from being declared and loaded only when the runtime inventory changes too. Peak process and system memory still require measurement. |
| Decoder-only Whisper GGML is an upstream-documented format | False. Upstream documents a Core ML encoder paired with a GGML model. The stripped container and compatible loader are local experiments. |
| Structural deletion requires filtering a PyTorch checkpoint before conversion | Too prescriptive. A checked byte-for-byte derivative of an authenticated GGML source is simpler for the current experiment and avoids reconversion drift. |
A caller must add @autoreleasepool or Core ML leaks about 5.8 MB per request |
Unsupported. The checked bridge already wraps prediction in @autoreleasepool; the exact leak rate has no local or primary-source proof. Measure repeated-call growth before changing ownership. |
phys_footprint alone predicts a macOS jetsam kill |
Unsupported as a singular rule. It is a valuable task ledger, but desktop termination and product pressure also depend on working set, wired and clean pages, compression, swap, other processes, and system state. |
MLComputeUnits.all proves ANE execution |
False. It makes CPU, GPU, and Neural Engine available. Operation placement needs MLComputePlan, Instruments, or equivalent physical evidence. |
| Last-row logits and shortened K/V caches are mandatory for decoder-only correctness | False. They are separate mutable-state optimizations. First prove structural weight deletion with the baseline state and public API unchanged. |
| Q5 decoder output should exactly match an FP16 decoder | Wrong control. Structural deletion should exactly match the complete model with the same quantization. Q5-versus-FP16 quality is a separate numerical gate. |
2. Architecture and Ownership
The target runtime has one acoustic path and one text path:
PCM
|
v
mel filters retained in GGML prefix
|
v
external Core ML encoder -----> [audio positions, audio state]
|
v
cross-attention K/V
|
v
retained GGML/Metal decoder
|
v
tokens and text
The Core ML package replaces native encoder compute. The decoder still needs encoder hyperparameters because they define the external output and cross-attention contract. Deleting native weights does not authorize changing those dimensions.
2.1. Keep the ABI-bearing prefix
A decoder-only derivative should preserve, byte-for-byte:
- file magic and model hyperparameters;
n_audio_ctx,n_audio_state,n_audio_head, andn_audio_layer;- text dimensions and quantization metadata;
- mel-filter dimensions and values;
- vocabulary bytes and token IDs; and
- every retained decoder record header and payload.
Preserving n_audio_layer may look odd after encoder deletion. In this legacy
container it also participates in model-family identification and expected
architecture. The runtime patch, not an invented header value, should make
native encoder tensors unreachable.
2.2. Delete only native encoder records
Delete the records named by the pinned encoder inventory:
- positional embedding;
- two convolution weights and biases;
- final encoder layer normalization; and
- every tensor in every native encoder transformer block.
Retain all decoder self-attention and cross-attention tensors. Cross-attention is part of the decoder even though its keys and values are derived from the external encoder output.
2.3. One owner per resource
A memory-minimal serialized product starts with:
- one model context owning decoder weights;
- one state owning mutable caches and schedulers;
- one Core ML model handle associated with that state; and
- one request at a time through the context.
The public whisper.cpp contract does not make one context concurrently safe for arbitrary state use. Add concurrency only after proving it is a product requirement and measuring the incremental state and framework residency.
3. Deriving the Artifact
3.1. Prefer exact record filtering over reconversion
For a pinned, already-authenticated GGML source, the simplest inspected derivative is a strict parser and copier:
- parse the prefix through vocabulary exact EOF boundaries;
- parse every tensor record with checked dimensions, type, name, byte count, and offset arithmetic;
- classify each name against frozen encoder and decoder inventories;
- copy the prefix and decoder record bytes without numerical conversion;
- reject unknown, duplicate, missing, malformed, or trailing records; and
- reparse the emitted artifact independently before sealing it.
This design isolates structural deletion. Filtering a PyTorch state dictionary and rerunning conversion can also work, but introduces converter, source-model, and quantizer drift that the deletion experiment does not need.
3.2. Seal both what remains and what disappeared
The derivation receipt should bind:
- complete source-model SHA-256 and byte size;
- prefix SHA-256 and byte size;
- ordered source tensor inventory;
- ordered retained and removed inventories;
- each retained record's source offset, byte length, type, shape, and hash;
- output SHA-256 and exact EOF;
- total retained and removed payload and physical record bytes;
- derivation script and focused test hashes; and
- the compatible runtime source, patch, build options, and binary closure.
An output hash alone cannot prove that the right tensors were retained. An inventory alone cannot prove payload identity. Bind both.
3.3. Current 8e2d arithmetic is local experimental evidence
The current campaign's authenticated Q5 physics shell is 537,819,875 bytes. An exact parse identifies 487 native encoder records and 52 decoder records. Preserving its 611,057-byte prefix and all decoder records projects a stripped artifact of 84,778,619 bytes, removing 453,041,256 physical file bytes.
Those numbers are local binary-format evidence, not upstream guarantees, resident-memory measurements, quality evidence, or a shipping verdict. The campaign note owns their current receipts and status.
4. Making the Loader Decoder-Only
The pinned loader currently calculates tensor metadata, declares all native encoder and decoder tensors, allocates backend buffers for them, and then requires the number of loaded records to equal the declared tensor count. Merely truncating the file therefore fails with missing tensors.
4.1. Change the expected inventory, not error tolerance
In an explicit external-encoder-only build or mode:
- do not create native encoder tensor metadata;
- do not add native encoder names to
model.tensors; - resize and create only the decoder-layer structures needed by text and cross-attention graphs;
- retain the stock unknown-tensor, shape, size, and loaded-count checks for the smaller inventory, then add exact type, duplicate, missing-name, and EOF checks; and
- log that native encoder capability is absent.
Do not make the generic loader accept arbitrary missing tensors. That converts a precise derivative into a partially initialized model that can fail far from the load boundary.
4.2. Make native encoder execution unreachable
Encoder pointers may remain null only if every path that can build or execute the native encoder is structurally unavailable. The contract should require:
- Core ML compiled in;
- native fallback compiled out or rejected for this artifact class;
- Core ML sidecar load before a usable state is returned;
- an error return when prediction fails; and
- no attempt to build convolution or encoder graphs with missing tensors.
The model-only _no_state initializer may reasonably authenticate and load
decoder weights before a Core ML state exists. If so, document that boundary:
model-only load can succeed, while state creation and inference must fail when
the sidecar is absent. A negative control must exercise the exact public entry
point the product uses.
4.3. Validate the external encoder result
Do not treat a non-null MLMultiArray as sufficient. Check:
- prediction returned without error;
- output feature name is the expected one;
- element type is Float32 when the bridge copies Float32;
- rank and every dimension match the compiled contract;
- strides/layout match the copy path, or copy by checked strides; and
- element count exactly matches the destination allocation.
The external output shape and the preserved header's audio context must agree. A shorter or longer output is a contract failure, not a reason to silently truncate, pad stale memory, or invoke the absent native encoder.
4.4. The existing autorelease pool is not the missing feature
The checked Objective-C++ bridge already places the Core ML prediction call in
an @autoreleasepool. Under ARC, locally owned objects also leave scope with
the function. That does not prove the entire application is leak-free, but it
invalidates the raw report's claim that this bridge lacks a pool and leaks a
fixed 5.8 MB on every request.
If repeated inference grows memory, distinguish:
- one-time Core ML specialization or allocator growth that plateaus;
- reusable framework caching;
- product-layer autoreleased objects outside the bridge; and
- genuinely unbounded retained objects or buffers.
Use allocations and VM traces to identify the owner before adding another pool or changing bridge lifetime.
5. Mutable State Is a Separate Optimization
Structural deletion targets immutable native encoder weights. Self-K/V, cross-K/V, pad caches, schedulers, compute buffers, token vectors, and logits remain. Their bounds depend on the product contract and should be changed only after same-quantization parity passes.
5.1. Logits
The pinned state reserves capacity proportional to vocabulary size times text context. A greedy product may ultimately need only current-step logits, but a last-row-only change can alter public accessor assumptions, batching behavior, and downstream pointer lengths. Measure committed memory first, then introduce an explicit API contract and regression tests if the saving is material.
5.2. Text and audio caches
Bounding text length can reduce mutable state, but every caller must reject an oversized request before graph construction. Bounding audio positions is not valid merely because typical input is shorter: the external encoder output, header metadata, cross cache, and decoder graph must all share the same frozen maximum.
For the current full-span 8e2d experiment, preserve the 1,500-position external encoder contract. Short-context state is a different artifact and quality experiment.
5.3. Disable unused product features explicitly
Greedy, no-timestamp dictation can avoid beam candidates and optional DTW state when the public product contract does not consume them. Record each disabled feature in the receipt. Do not infer that a zero-valued request flag prevented an earlier initialization allocation; inspect the construction order.
6. Proving Memory Reduction on macOS
Artifact bytes, backend allocation logs, process footprint, and system pressure answer different questions. No single column proves the whole result.
| Evidence | What it establishes | What it does not establish |
|---|---|---|
| Artifact size and inventory | Native encoder records are absent on disk | Runtime allocation or Core ML residency |
| Backend weight-buffer log | Declared decoder tensors occupy a smaller backend buffer | Peak task or system memory |
TASK_VM_INFO.phys_footprint and peak ledger |
Memory charged to the worker task at sample times | Complete cross-process residency or universal termination risk |
/usr/bin/time -lp maximum RSS and footprint |
Coarse process peaks over a complete invocation | Exact phase attribution |
footprint and vmmap -summary |
Snapshot categories and VM-region attribution | A missed earlier peak |
| Core ML service census | Correlated work outside the app process | Causal attribution without a matched control |
vm_stat, swap, and memory-pressure series |
Host-level pressure during the experiment | Per-model ownership |
MLComputePlan or Instruments placement |
Eligible device use by operation on the tested build | Future devices, OS versions, or artifacts |
Sample at named phases:
- before model load;
- after decoder weights load;
- after state and Core ML sidecar initialization;
- after the first prediction and specialization;
- across warmed repeated predictions; and
- after state/context release and an idle observation window.
Compare complete-model and stripped-model processes with the same backend, Core ML tree, input, request, thread count, warmup, host admission, and sample cadence. Do not overlap the old and new contexts during replacement.
6.1. ComputeUnits.all is eligibility, not placement proof
Core ML Tools documents ComputeUnits.ALL as allowing all available units:
CPU, GPU, and Neural Engine. The same documentation exposes
MLComputePlan.get_compute_device_usage_for_mlprogram_operation to inspect
planned operation use. Therefore a successful load or prediction under ALL
does not, by itself, prove ANE residency.
For a physical claim, bind the compiled-tree hash and capture an operation-level compute plan or Instruments trace on the exact target host. Report boundary CPU/GPU operations rather than collapsing a mixed graph into “runs on ANE.”
7. Correctness and Negative Controls
The structural-deletion control must use the same quantized decoder and same external Core ML encoder on both sides. Require:
- identical token IDs in order;
- identical EOT presence and stop reason;
- identical decoded-text bytes;
- identical token ceilings and request options;
- deterministic output across repetitions; and
- no latency regression beyond a preregistered tolerance.
Include full-span inputs near both 15 and 30 seconds. A load-only pass cannot prove that null native-encoder pointers stay unreachable during real prediction.
Named negative controls should reject:
- missing or renamed Core ML sidecar;
- Core ML prediction error or wrong output shape/type;
- one missing decoder record;
- one duplicate decoder record;
- an unknown encoder record reintroduced into the stripped artifact;
- wrong tensor type, dimensions, or payload length;
- truncated input and trailing bytes;
- native-fallback configuration; and
- more tokens or audio positions than the bounded state allows.
Q5-versus-FP16 WER and pathology qualification remains separate. Exact stripped-Q5 versus complete-Q5 parity proves deletion is lossless; it does not prove Q5 is good enough for production.
8. Failure Modes and Debugging Playbook
8.1. Loader reports an unknown or missing tensor
The derivative inventory and runtime declaration disagree. Print the first unexpected name, dimensions, type, record offset, and expected inventory hash. Do not suppress the error or permit partial loading.
8.2. Load succeeds, then native graph construction crashes
A fallback or graph-builder path still reaches a deleted encoder pointer. Reproduce with the Core ML sidecar missing, then make the external-encoder-only mode fail before any inference graph can be built.
8.3. Output repeats or hallucinates after the split
First compare the external encoder output shape and bytes, cross-cache construction, request reset, and exact decoder artifact against the complete control. Do not attribute the symptom to encoder deletion until those contracts match. Stale cross-K/V and audio-context mismatch can produce plausible text without a clean crash.
8.4. The file shrinks but peak memory does not
Check the model-load buffer log. If it still reports the complete weight size, the runtime still declares native tensors. If that buffer shrank, split the remaining peak among Core ML initialization, caches, compute arenas, temporary read buffers, dynamic-library load, and overlapping contexts.
8.5. Memory grows across repeated predictions
Run enough iterations to distinguish lazy plateau from linear growth. Mark pool boundaries, model/state lifetime, and request completion in the trace. Because the bridge already has a prediction-local autorelease pool, inspect actual retained owners before diagnosing an autorelease leak.
8.6. Core ML loads, but ANE evidence is absent
This is expected under a permissive compute-unit request. Capture the exact compiled model's compute plan or Instruments trace. Treat load logs as Core ML utilization evidence only, not device-placement evidence.
9. Anti-Patterns
| Anti-pattern | Why it fails | Do this instead |
|---|---|---|
| Delete bytes and relax “not all tensors loaded” | Accepts partially initialized arbitrary models | Declare an exact smaller inventory and retain strict EOF/count checks |
| Change header audio dimensions to describe missing weights | Breaks external output and cross-attention sizing | Preserve ABI-bearing metadata; remove native capability in code |
| Keep native fallback “for safety” | Deleted tensors make fallback impossible | Fail state creation or prediction closed |
| Reconvert and requantize during the deletion control | Mixes structural and numerical changes | Copy authenticated decoder records byte-for-byte |
| Claim the disk-byte delta as RAM saved | Ignores state, services, allocator behavior, and peaks | Run matched process and system measurements |
| Add another autorelease pool from a fixed leak anecdote | Changes lifetime without identifying an owner | Prove plateau versus leak with traces |
Claim ANE from ComputeUnits.all or a load log |
Availability and Core ML use do not identify placement | Bind MLComputePlan or Instruments evidence |
| Shrink caches and logits in the first patch | Makes a lossless parity failure ambiguous | Prove weight deletion first, optimize state second |
| Compare stripped Q5 directly with FP16 for exactness | Quantization itself can change tokens | Compare stripped and complete models at the same quantization |
10. Decision Framework
| Approach | Value | Main risk | Use when |
|---|---|---|---|
| Stock hybrid model | Upstream-documented control | Retains unused native encoder allocation | Establishing correctness and memory baseline |
| Full-file validate-and-skip loader | Isolates allocation deletion without changing file bytes | More runtime parsing/skip machinery; forbidden bytes remain shippable | Temporary losslessness falsifier |
| Authenticated stripped artifact plus explicit loader | Smallest disk and declared weight inventory | Local format/runtime fork needs strong receipts | Preferred experiment after mechanics are understood |
| Reconverted decoder-only checkpoint | Can integrate into a future producer | Adds conversion and quantization drift | Only when source-level production generation is required |
| Mutable-state reduction | Additional RAM after weights | API, capacity, and correctness regressions | After structural parity and measured state attribution |
| Core ML compression | May reduce package and residency | Numerical drift and placement changes | Last, under a separate physical qualification |
The clean promotion order is:
- complete-model same-quantization control;
- stripped artifact and fail-closed loader;
- exact inference parity and malformed-input controls;
- matched local physical-memory proof;
- exact target-Mac latency, placement, memory, and swap proof;
- mutable-state experiments one mechanism at a time; and
- independent decoder quantization and product-quality qualification.
Works Cited
- whisper.cpp pinned source tree — source pin used for the legacy loader, tensor, state, and Core ML boundary.
- whisper.cpp Core ML support — upstream hybrid encoder setup and runtime load-log contract.
- Core ML Tools compute-unit options —
ALL, CPU, GPU, and Neural Engine eligibility choices. - Core ML Tools model utilities —
MLComputePlanoperation-device and estimated-cost inspection. - Low-memory Whisper Turbo inference — locally corrected Darwin accounting, pinned allocation topology, and memory receipt design.
- Physical Core ML plus whisper.cpp verification — local fail-closed and placement-evidence contract.