# Low-Memory Whisper Turbo Inference on Apple Silicon

> **Collective Library edition.** This is the complete technical report. Private filesystem paths, internal run identifiers, campaign-control notes, and repository navigation were removed. Technical claims, code, measurements, evidence labels, citations, corrections, and falsification criteria are preserved.

August 12, 2026

## Executive Summary

The minimum-memory architecture is not “mmap the model and tune Darwin page
hints.” The pinned whisper.cpp loader does not mmap the model at all. It opens
the custom GGML `.bin` with `std::ifstream`, allocates backend buffers for every
declared encoder and decoder tensor, and streams the weight bytes into those
buffers. When Core ML is enabled, a separate `MLModel` performs encoder
inference, but the native encoder weights remain allocated and loaded.

That establishes a deletion-first ladder:

1. instrument the current resident FP16 hybrid and seal its real process and
   system footprint;
2. run the stock whole-model Q5_0 control to learn the safe quantization and
   bandwidth effect;
3. add an explicit fail-closed external-encoder load mode that never allocates
   or loads native encoder tensors;
4. quantize only the remaining decoder-side weights after the lossless
   encoder-deletion control proves exact output parity;
5. right-size one serialized state for the product's bounded audio and search
   contract; and
6. compress the Core ML encoder only after the host-side waste is gone.

`TASK_VM_INFO.phys_footprint` is an important app-process metric, but it is not
the sole truth and macOS has no documented universal “jetsam threshold” for a
desktop process. A product receipt must pair the worker's footprint with RSS,
dirty/clean/wired VM categories, relevant service processes, system
compression and swap deltas, latency, and transcript parity.

The raw report's mandatory caller-side `@autoreleasepool` and implied SIGBUS
cure are also unsupported for this pin. The pinned Metal backend already wraps
graph execution and tensor-transfer paths in autorelease pools, retains and
releases command buffers explicitly, and the Core ML wrapper uses ARC plus a
prediction-local pool. Add another pool only when a measured application-level
Objective-C loop accumulates autoreleased objects; do not treat it as a generic
Metal crash fix.

## 1. Corrected Runtime Facts

### 1.1. Claims from the raw report that require correction

| Raw-report claim | Corrected verdict |
| --- | --- |
| Whisper weights are loaded through GGUF/GGML mmap | **False for the pinned runtime.** Whisper uses its custom sequential GGML `.bin`, opens it with `std::ifstream`, allocates backend tensors, and reads bytes into them. |
| Clean mapped pages make model weights nearly free in `phys_footprint` | **Not applicable to the pinned loader.** The weight buffers are allocated runtime storage, not a lazy read-only file mapping. |
| `phys_footprint` is the only metric that predicts termination | **Overstated.** It is the task's charged footprint ledger and a primary metric for dirty memory impact. Clean working set, wired memory, opaque service allocations, compressor activity, swap, and system pressure still matter. |
| Core ML creates exactly two physical copies of the encoder | **Not established.** The source proves a full native encoder allocation plus a separate Core ML model instance. Core ML's physical representation and accounting require measurement. |
| `mlock` converts clean file pages into charged footprint | **Imprecise and irrelevant here.** `mlock` keeps mapped pages resident and is subject to wired and resource limits; the pinned model is not file-mapped. |
| `MADV_WILLNEED` is the recommended warm-model solution | **Not applicable here.** `madvise` only advises an existing mapped range. The pinned model weights live in backend allocations. |
| `MADV_FREE` safely leaves model pages available for reuse | **Dangerous.** `MADV_FREE` says the application no longer needs the contents and permits immediate reuse. It is not a model-weight cache hint. |
| Every inference loop needs a caller `@autoreleasepool` to avoid SIGBUS | **Unsupported.** The pinned Metal and Core ML paths already contain scoped pools. SIGBUS has many causes; no primary source ties it to a missing outer pool here. |
| One context can serve parallel states with zero duplication | **Contradicted by the public header.** The C interface says the same `whisper_context` must not be used concurrently. Every state also owns caches, schedulers, backends, and a Core ML model handle. |
| Setting `audio_ctx=600` automatically shrinks cross-K/V allocation | **False for this pin.** State initialization allocates cross and pad caches from the model's full `n_audio_ctx`, before request parameters set `exp_n_audio_ctx`. |
| Q5_0 has negligible accuracy loss and is the Apple Pareto frontier | **Unverified for this product.** Upstream supports Q5_0; accuracy and latency remain model-, corpus-, and device-specific promotion gates. |

### 1.2. The model file is not GGUF

The pinned Whisper loader expects its own ordered GGML binary. It checks
`GGML_FILE_MAGIC`, then reads hyperparameters, mel filters, vocabulary, and
named tensors. Vendored GGUF support elsewhere in the implementationsitory does not make
this load path GGUF.

This distinction matters for memory engineering. Advice written for a GGUF
mmap loader cannot be transferred to this Whisper runtime without inspecting
the actual loader.

### 1.3. What Core ML replaces—and what it does not

Core ML replaces encoder **compute** after state initialization:

```text
audio
  |
  v
mel preparation in whisper.cpp
  |
  v
Core ML encoder MLModel --------> encoder embeddings
                                      |
                                      v
                          whisper.cpp cross cache
                                      |
                                      v
                          whisper.cpp decoder
```

It does not currently replace native encoder **storage**. The model loader
unconditionally constructs encoder tensor metadata, allocates all corresponding
backend buffers, and loads every tensor before `whisper_init_state` opens the
Core ML sidecar.

The source therefore proves structural redundancy: the native encoder tensor
payload exists even though the native encoder graph is skipped. It does not
prove that Core ML adds a byte-for-byte second FP16 physical copy. Core ML may
compile, compress, cache, share, or charge its representation differently, and
some allocations may reside in framework or service processes.

## 2. Allocation Topology

### 2.1. Model-global context

`whisper_context` owns:

- model hyperparameters, vocabulary, and mel filters;
- tensor metadata;
- backend buffers containing all loaded model weights; and
- optionally a default `whisper_state` when a non-`_no_state` initializer is
  used.

The pinned loader's decisive sequence is:

1. create every encoder and decoder tensor description;
2. allocate backend buffers for those descriptions;
3. read every weight tensor directly into a host-capable buffer, or through a
   temporary buffer into another backend; and
4. reject unknown, missing, or shape-mismatched tensors.

The file is closed after load. Deleting or advising file-cache pages cannot
release the model's allocated backend storage.

### 2.2. Per-state allocations

Every `whisper_state` owns substantial mutable state:

- backend instances;
- self-attention, cross-attention, and flash-attention pad caches;
- convolution, encoder, cross-attention, and decoder scheduler storage;
- mel, logits, batch, candidate, and result vectors;
- optional DTW and VAD structures; and
- under `WHISPER_USE_COREML`, a separately initialized Core ML model handle.

When an external encoder is live, the pin omits the native encoder scheduler,
but it still allocates the convolution shim, cross-attention scheduler,
decoder scheduler, all caches, and all model weights.

Use `whisper_init_from_file_with_params_no_state` followed by exactly one
`whisper_init_state` when custom state ownership is required. Calling the
ordinary initializer and then allocating an additional state silently creates
two states.

### 2.3. One serialized state is the memory floor

The public header's contract says the same context must not be used by multiple
threads concurrently. `whisper_full` is explicitly not thread-safe for the same
context. Separate states expose finer-grained APIs, but the header does not
promote concurrent same-context execution as supported.

For a dictation product that finalizes one active speaker at a time, the
minimum-memory topology is therefore simple:

```text
one model context
        |
        +-- one state
                |
                +-- one serialized inference queue
```

Do not create a state pool until product concurrency requires it. Each state
adds caches and schedulers and initializes another Core ML model object.
Framework-internal sharing may reduce the incremental cost, but that is a
measurement question, not a zero-duplication guarantee.

### 2.4. Metal shared buffers do not imply a CPU-plus-GPU copy

The pinned Metal backend can allocate CPU storage and wrap it with
`newBufferWithBytesNoCopy` using shared storage. In that path the `MTLBuffer`
references the same allocation; it is not evidence of a second full GPU copy.
The backend keeps the underlying allocation alive with the Metal buffer and
releases them together.

Other paths can use private Metal storage and explicit blits. Record the actual
backend buffer names and storage path from the pinned build. “Unified memory”
does not mean every API object aliases the same bytes, and “Metal buffer” does
not by itself mean duplicate DRAM.

## 3. Exact State Arithmetic for Turbo

### 3.1. Architecture constants

The official `openai/whisper-large-v3-turbo` configuration and pinned GGML
header use:

- vocabulary: 51,866;
- audio context: 1,500 positions;
- model width: 1,280;
- encoder layers: 32;
- text context: 448 tokens; and
- decoder layers: 4.

The model's FP16 file and loader-reported weight size are recorded in the
local campaign note.
File size is artifact evidence; it is not a peak-memory measurement.

### 3.2. K/V allocations use padded maxima

The pinned state initializer uses FP16 intermediate caches and rounds context
lengths to multiples of 256. For Turbo:

| Cache | Source dimensions | Allocated bytes |
| --- | --- | ---: |
| Self K and V | `4 layers * 512 text slots * 1280 width * 2 arrays * 2 bytes` | 10,485,760 |
| Cross K and V | `4 layers * 1536 audio slots * 1280 width * 2 arrays * 2 bytes` | 31,457,280 |
| Encoder pad K and V | `1 layer * 1536 audio slots * 1280 width * 2 arrays * 2 bytes` | 7,864,320 |
| **Total** | All three caches | **49,807,360** |

This corrects two common mistakes:

1. use the padded 512 and 1,536 allocations, not logical 448 and 1,500; and
2. include the separate pad cache.

Setting request `audio_ctx=600` changes graph views after state creation. It
does not change these allocations. A bounded-state patch would need to pass the
product's maximum audio positions into state initialization and allocate
`GGML_PAD(max_audio_ctx, 256)` deliberately.

That patch has a small ceiling compared with native encoder-weight deletion.
At 600 positions, the rounded audio cache is 768 positions, so shrinking cross
and pad caches can save only the difference between the 1,536- and 768-slot
allocations. Prove it after the gigabyte-scale waste is removed.

### 3.3. Reserved address space is not physical footprint

The pin reserves `n_vocab * n_text_ctx` floats for the state's logits vector:
92,943,872 bytes for Turbo. `std::vector::reserve` obtains capacity, but Darwin
does not necessarily charge every untouched demand-zero page to resident or
physical footprint immediately. Runtime `resize` and tensor copies determine
how much becomes touched.

Do not claim a 92.9 MB physical saving merely by removing the reserve. Sample
before and after model load, state creation, first prefill, steady greedy
decoding, and teardown. If the pages remain untouched, deleting the reserve
may reduce virtual size without materially reducing peak inference RAM.

### 3.4. Compute buffers are empirical

The runtime logs scheduler sizes for convolution, encoder, cross-attention, and
decoder paths. Those values depend on backend, flash attention, graph shape,
and build. Capture the log lines in the receipt instead of deriving them from
parameter count.

An external encoder should remove the native encoder scheduler. It does not
remove:

- the mel/conv shim;
- encoder output storage;
- cross-attention K/V construction;
- decoder compute storage; or
- native encoder weights in the current loader.

## 4. Darwin Memory Accounting

### 4.1. No single column is enough

The installed `footprint(1)` documentation defines a process footprint around
dirty memory charged to that process and recommends reducing dirty memory
first. It separately reports clean, reclaimable, swapped, and wired memory.
Clean file-backed working sets can still cause latency thrash; wired memory
cannot be paged, compressed, or reclaimed.

Use a metric set rather than a winner:

| Metric | What it answers | Important blind spot |
| --- | --- | --- |
| `TASK_VM_INFO.phys_footprint` | Current footprint charged to the worker task | Does not attribute opaque work in other processes or describe clean working-set latency |
| `ledger_phys_footprint_peak` | Peak charged footprint observed for the worker | Requires a sufficiently new `TASK_VM_INFO` revision; still worker-only |
| `resident_size` / RSS | Pages currently resident for the process | Includes categories with different reclaimability and can overstate durable pressure |
| `footprint` dirty/clean/reclaimable/wired columns | Category-level task or multi-process accounting | Cross-process inspection can require root; sampling has overhead |
| `vmmap -summary` | VM regions, dirty/swapped allocation, and mapped-file categories | Snapshot, not a peak; permissions may restrict inspection |
| `vm_stat` deltas | System compressor, page-in/out, and swap activity | System-wide and sensitive to unrelated work |
| `memory_pressure` snapshot | Current system memory statistics and free percentage | Not per-process; `-S` simulates pressure and invalidates a benchmark |
| Core ML/ANE service census | Whether work appears outside the worker | Process names and accounting can change by OS version; causality needs a control |

macOS desktop process termination is not specified as one public footprint
threshold. A low worker footprint can coexist with high system pressure, and a
high RSS can include reclaimable or clean pages. Report both.

### 4.2. In-process sampler

The installed SDK exposes current and peak fields through `TASK_VM_INFO`.
Sample in-process so the worker can bind memory points to exact inference
events:

```cpp
#include <mach/mach.h>
#include <mach/task_info.h>

#include <cstdint>

struct DarwinTaskMemory {
    uint64_t resident_bytes;
    uint64_t resident_peak_bytes;
    uint64_t compressed_bytes;
    uint64_t footprint_bytes;
    uint64_t footprint_peak_bytes;
};

static bool read_darwin_task_memory(DarwinTaskMemory & out) {
    task_vm_info_data_t info = {};
    mach_msg_type_number_t count = TASK_VM_INFO_COUNT;
    const kern_return_t result = task_info(
        mach_task_self(),
        TASK_VM_INFO,
        reinterpret_cast<task_info_t>(&info),
        &count);

    if (result != KERN_SUCCESS || count < TASK_VM_INFO_REV1_COUNT) {
        return false;
    }

    out.resident_bytes = info.resident_size;
    out.resident_peak_bytes = info.resident_size_peak;
    out.compressed_bytes = info.compressed;
    out.footprint_bytes = info.phys_footprint;
    out.footprint_peak_bytes =
        count >= TASK_VM_INFO_REV3_COUNT &&
                info.ledger_phys_footprint_peak > 0
            ? static_cast<uint64_t>(info.ledger_phys_footprint_peak)
            : 0;
    return true;
}
```

Treat a zero peak as unavailable unless the surrounding API call failed. Do
not silently replace it with the current value.

### 4.3. Required measurement phases

For every candidate arm, capture the same event sequence:

1. quiet-host admission and system baseline;
2. worker process started, before model load;
3. after native model context load;
4. after the single state and Core ML model load;
5. after first specialization or warmup;
6. immediately before each measured request;
7. at a high enough cadence during the request to observe the peak;
8. immediately after final text;
9. after an idle interval; and
10. after explicit state/context teardown.

A before/after snapshot is insufficient when an intermediate compute buffer is
the peak. Prefer in-process phase markers plus a 50-100 ms external sampler.
Record sampler overhead with a no-sampler latency control.

### 4.4. Process and system scope

At minimum, preserve:

```bash
# Per-process VM category snapshot.
vmmap -summary "$WORKER_PID"

# Current RSS in KiB; not a peak and not a footprint substitute.
ps -o pid=,rss=,vsz=,command= -p "$WORKER_PID"

# System totals followed by one-second deltas.
vm_stat 1

# Current system memory statistics. Do not add -S to a measurement run.
memory_pressure
```

Use `footprint --sample` when privileges and overhead are acceptable. When
comparing multiple related processes, use its multi-process de-duplication
rather than summing RSS columns blindly.

Core ML may perform work in framework or service processes that are not child
processes of the app. Record a process census before and during the Core ML arm
and pair it with a native-encoder control. Do not hard-code a service name as a
stable API or attribute every system-service delta to the model.

## 5. The Deletion-First Optimization Ladder

### 5.1. Stage 0: seal the existing FP16 hybrid

Before changing allocation, measure the exact current product candidate. The
receipt must bind:

- model file SHA-256 and byte size;
- Core ML package and compiled-tree identity;
- whisper.cpp commit and dirty-source state;
- build flags, linked libraries, and compute-unit policy;
- process and system memory series;
- runtime allocation log lines;
- p50, p95, and maximum stop-to-final latency; and
- exact tokens, EOT, final text, and quality-slice results.

This stage converts “approximately 2-3 GB” into a physical baseline. Without
it, later deltas cannot be attributed.

### 5.2. Stage 1: stock whole-model Q5_0

The pinned quantizer supports Q5_0 and applies it to eligible two-dimensional
tensors. It deliberately skips encoder convolution biases, encoder positional
embeddings, and decoder positional embeddings; non-two-dimensional tensors are
also retained.

Q5_0 is a high-leverage first control because the loader stores weights in
their quantized tensor type. It can reduce the large native backend buffer and
memory bandwidth without a custom runtime patch.

It is not lossless. Require:

- artifact/header validation and exact source binding;
- model-load success with the same Core ML sidecar;
- latency and memory measured with the same protocol;
- exact-output diagnostics on frozen rows; and
- full WER and product-pathology qualification before promotion.

Do not state an expected byte size or accuracy delta before the generated
artifact and bakeoff exist.

### 5.3. Stage 2: skip native encoder allocation under fail-closed Core ML

This is the largest source-verified deletion opportunity. The lossless control
should retain:

- the original header, vocabulary, filters, and decoder tensors;
- decoder embeddings and output projection;
- all decoder self- and cross-attention weights;
- the external encoder output ABI; and
- a hard requirement that the exact Core ML model loads successfully.

It should omit backend allocation for:

- encoder positional embeddings;
- both encoder convolution layers; and
- all 32 native encoder transformer layers.

Do not begin by inventing a new “decoder-only GGML” container. The smaller
first patch can keep the existing full artifact and teach an explicit
external-encoder-only loader mode to validate encoder tensor names and shapes
while seeking or discarding their bytes without allocating backend tensors.
That isolates RAM deletion from file-format surgery.

The current generic loader callback has `read`, `eof`, and `close`, but no
seek/skip callback. Reading a skipped giant tensor into a temporary vector
would recreate the peak being removed. Add either:

1. a bounded chunk discard path; or
2. an optional checked `skip` callback implemented by the file loader.

The `skip` path must reject truncated input, overflow, unexpected tensor names,
wrong shapes, wrong types, missing tensors, duplicate tensors, and a Core ML
load failure. It must never silently fall back to native encoding because the
native weights do not exist.

The external encoder call must fail closed too. A successful model load is not
enough: propagate `predictionFromFeatures` errors, require exactly one Float32
`[1, 600, 1280]` output, validate its element count and strides before copying,
and return an encoding failure rather than decoding stale or malformed state.
An unchecked `memcpy(output.count)` is incompatible with a runtime whose native
encoder fallback has deliberately been deleted.

Promotion gate: exact token IDs, EOT, text, logits within a preregistered
tolerance if captured, and latency parity against the full FP16 hybrid. The
memory delta is then causally attributable to unused native encoder storage.

### 5.4. Stage 3: quantize the retained decoder

Only after Stage 2 proves lossless deletion should quantization be applied to
the remaining eligible decoder tensors. This cleanly separates:

- exact-output changes caused by deleting unused tensors—which should be none;
  from
- numerical changes caused by decoder quantization.

A compact decoder-only shipping artifact can follow later if disk size and
load I/O justify a new format. The runtime patch and its validation rules are
the hard part; truncating bytes first merely creates an incompatible file.

### 5.5. Stage 4: right-size mutable state

After weight waste is gone, test smaller allocations in descending value:

1. exactly one state and one serialized request;
2. greedy search with one decoder rather than beam candidates;
3. DTW/token-timestamp structures disabled when the product does not consume
   them;
4. bounded cross and pad caches created from the product's maximum external
   encoder positions; and
5. logits capacity and decoder work buffers sized from observed maximum batch
   rather than theoretical text context.

Every bound needs a fail-closed input check. A 600-position state must reject
an encoder output longer than 600 rather than overrun, silently truncate, or
reuse stale cache tails.

### 5.6. Stage 5: compress Core ML last

Core ML palettization can reduce package storage and may reduce runtime weight
cost, but package bytes are not resident bytes. The compiler can transform or
expand representations, and compute placement can change.

For each compression arm, separately prove:

- portable package size and compiled-tree size;
- numerical parity against the dense encoder;
- end-to-end transcript and WER behavior;
- actual ANE placement;
- cold specialization and warm-load behavior;
- app and service-process memory; and
- stop-to-final latency under the same contention contract.

Do not combine decoder quantization, native-encoder deletion, and Core ML
compression in the first experiment. One changed mechanism per arm preserves
causal evidence.

### 5.7. Idle residency is a product policy, not a kernel trick

Loading while the user speaks can hide resident initialization from
stop-to-final latency. Releasing the context after an idle timeout can lower
idle RAM. Neither reduces peak inference memory.

Measure three policies:

| Policy | User benefit | Memory cost | Risk |
| --- | --- | --- | --- |
| Always resident | Lowest request-start variance | Highest idle RAM | System pressure while unused |
| Load on recording start | Hides load behind speech | Peak unchanged, idle lower | Short utterances may stop before load finishes |
| Idle timeout | Amortizes nearby requests | Tunable | First request after timeout pays load/specialization |

Core ML's first device specialization is a separate product event. Do not
confuse install-time compilation, warm model load, and steady prediction.

## 6. `mmap`, `mlock`, and `madvise`

### 6.1. Why they are not current levers

The pinned Whisper path does not expose a mapped model-weight range. Therefore:

- there is no model mapping to lock;
- there is no model mapping to prefetch with `MADV_WILLNEED`;
- dropping the source file's cache does not free backend weights; and
- a benchmark of file-cache advice would test I/O, not resident inference RAM.

Delete this work from the immediate plan.

### 6.2. If a future loader introduces mmap

Treat these APIs according to their platform contracts:

- `mlock` keeps the addressed pages resident until unlocked and is bounded by
  process and system limits. It increases non-reclaimable pressure and should
  be an explicit latency-versus-memory experiment, not a default.
- `MADV_WILLNEED` is advisory. The man page does not promise asynchronous
  completion, full prefetch, or future residency.
- `MADV_DONTNEED` says the range is not expected soon; it is not a durability
  or latency guarantee.
- `MADV_FREE` permits the VM to reuse contents. Never apply it to live model
  weights.

An mmap rewrite might reduce dirty weight storage by relying on clean
file-backed pages, but Metal compatibility, tensor alignment, quantized access,
page-fault tails, and Core ML coexistence all need a separate design and
physical falsifier. It is not prerequisite to removing unused encoder tensors.

## 7. Autorelease Pools, No-Copy Buffers, and SIGBUS

### 7.1. What the pin already does

The pinned sources establish:

- `src/coreml/whisper-encoder.mm` refuses to compile without ARC;
- Core ML prediction is wrapped in `@autoreleasepool`;
- Metal graph compute and tensor transfer paths contain their own
  `@autoreleasepool` blocks;
- command buffers created with unretained references are explicitly retained,
  replaced, synchronized, and released by the backend; and
- shared Metal buffers can wrap backend-owned CPU storage with a nil
  deallocator while that storage remains owned by the backend buffer.

That is not evidence of a pool leak.

### 7.2. What an outer pool can and cannot do

An application-level `@autoreleasepool` around a long Objective-C loop can
bound objects created by the application or frameworks and released by
autorelease. Use it when an allocation trace shows growth between pool drains.

It cannot repair:

- a CPU pointer freed before an asynchronous GPU operation completes;
- out-of-bounds tensor views or incorrect strides;
- a command buffer or resource that is explicitly retained and never released;
- invalid model bytes;
- an OS or driver fault; or
- memory pressure caused by live, required buffers.

### 7.3. No-copy lifetime rule

Apple's no-copy Metal initializer wraps caller-provided bytes. The caller must
keep those bytes valid for the lifetime and use of the `MTLBuffer`. In an
asynchronous path, lifetime must extend through GPU completion, not merely the
encoding call.

When SIGBUS occurs, preserve the crash report, fault address, GPU error state,
command-buffer status, tensor name and range, buffer ownership, and completion
ordering. Do not diagnose “missing autorelease pool” from the signal alone.

## 8. Decision Framework

| Mechanism | Expected memory target | Accuracy risk | Latency risk | Evidence status |
| --- | --- | --- | --- | --- |
| Measure current FP16 hybrid | None; establishes truth | None | Sampling overhead | Required first |
| Whole-model Q5_0 | Native backend weights | Real; must qualify | May improve bandwidth or add dequant cost | Source-supported candidate |
| Skip native encoder allocation | Unused native encoder weights | None if external path is exact and fail-closed | Load path changes; inference should match | Highest-value hypothesis from source |
| Quantize retained decoder | Decoder weights and bandwidth | Real; must qualify | Device-dependent | Test after lossless deletion |
| Single `_no_state` context plus one state | Duplicate caches, schedulers, Core ML objects | None for serialized product | Removes concurrency | Source-verified topology |
| Bound cross/pad cache to 600 | Tens of MB, not GB | Input-bound failure if contract violated | Small possible gain | Secondary source-supported patch |
| Remove oversized logits reserve | Virtual capacity; physical saving uncertain | Possible batch regression | Probably small | Measure before implementation |
| Disable DTW and beam search | Optional mutable state | Product-feature and decode-quality tradeoff | Usually favorable | Product-contract dependent |
| LUT-compress Core ML encoder | Core ML package and possibly runtime weights | Numerical and placement risk | Compile and prediction risk | Last-stage physical experiment |
| Idle unload | Idle memory only | None | Reload/specialization tail | Product policy |
| Add model mmap | Dirty weight storage, possibly | Loader/backend rewrite risk | Page-fault tails | Separate future program |
| Add `mlock`/`madvise` now | None in current loader | None | Engineering distraction | Delete |

## 9. Failure Modes and Debugging Playbook

### 9.1. File size falls but peak footprint does not

Likely causes:

- the quantized file still expands into another backend type;
- Core ML or compute buffers dominate the peak;
- the sampler missed the true peak;
- both old and new contexts overlap during model replacement; or
- allocator high-water pages remain charged after a transient conversion.

Check backend buffer log sizes, phase markers, context overlap, and
`footprint`/`vmmap` categories before changing quantization again.

### 9.2. Worker footprint falls but system pressure does not

Run the native-encoder and Core ML controls with identical host admission.
Inspect relevant service-process and system deltas. Core ML work may be charged
outside the worker, but unrelated services can move at the same time. Require
repeatable paired deltas before attribution.

### 9.3. A second state causes a large jump

Verify the initializer. If the ordinary context initializer created a default
state and the application then called `whisper_init_state`, free the duplicate
and switch to the `_no_state` initializer. Next compare state creation with and
without Core ML to separate caches/schedulers from framework model residency.

### 9.4. `audio_ctx=600` did not shrink initialization logs

This is expected for the pin. The request parameter is assigned after state
creation, while K/V caches were allocated from full model hyperparameters.
Implement a checked state-init bound or stop claiming a memory saving.

### 9.5. RSS is much larger than footprint

Use `footprint` and `vmmap` to split dirty, clean, reclaimable, wired, and
shared categories. Do not optimize the numerical difference itself. Optimize
the category that limits the product: charged pressure, clean working-set
thrash, wired allocations, or latency after eviction.

### 9.6. Memory grows across repeated requests

Separate three patterns:

1. one-time lazy allocation that plateaus;
2. allocator caching that remains reusable; and
3. unbounded growth proportional to request count.

Run enough iterations to distinguish them, add explicit pool-drain markers at
the application boundary, and record live object/VM categories. A stable higher
plateau is not automatically a leak; linear growth is not automatically an
autorelease problem.

### 9.7. Q5 changes one proper noun or number

That is an accuracy failure for exact-output parity even if aggregate WER barely
moves. Keep the full pathology matrix: names, numbers, negations, punctuation,
false starts, repetitions, silence, and clean dictation. Promote using the
product's quality policy, not file size or average WER alone.

### 9.8. Metal faults after a no-copy change

Audit ownership and ordering first:

1. identify the exact pointer and byte range;
2. prove alignment and buffer bounds;
3. prove the CPU allocation remains alive;
4. wait for or otherwise order GPU completion before reuse/free;
5. inspect command-buffer error details; and
6. rerun with a copied-buffer control.

An outer autorelease pool is not a substitute for those proofs.

## 10. Production Qualification Receipt

Every memory candidate should emit one sealed record with:

### Artifact identity

- source model and revision;
- GGML model SHA-256, byte size, and ftype;
- Core ML package and compiled-tree digests;
- whisper.cpp commit, patch digest, and dirty status;
- executable and linked-runtime digests; and
- OS, hardware model, RAM capacity, and boot identity.

### Runtime topology

- context initializer used;
- state count;
- Core ML model count;
- backend buffer names and logged sizes;
- compute-unit setting;
- flash-attention, timestamp, VAD, search, and audio-context settings; and
- serialized or concurrent request policy.

### Memory series

- worker current and peak `phys_footprint`;
- worker current and peak resident size;
- worker compressed bytes;
- `footprint` category summaries when available;
- relevant service-process series or an explicit “not attributable” marker;
- system wired/compressor/page-in/page-out/swap deltas; and
- baseline, post-load, post-state, warmup, per-request peak, idle, and teardown
  phase timestamps.

### Product outcomes

- stop-to-final p50, p95, and maximum;
- cold load and first-specialization time reported separately;
- exact token IDs, EOT, and final text for diagnostic fixtures;
- WER and pathology gates on frozen slices;
- silence false alarms;
- memory growth slope over a long repeated-request soak; and
- cancellation, fallback, and idle-unload behavior.

The promotion rule is conjunctive: lower memory is useful only if latency,
quality, stability, and fail-closed behavior still pass.

## 11. Anti-Patterns

| Anti-pattern | Why it fails | Use instead |
| --- | --- | --- |
| Quote model file size as RAM | Ignores allocator, runtime, state, services, and compression | Phase-aligned physical measurements |
| Treat RSS as the sole metric | Mixes categories with different reclaimability | Footprint plus VM categories and system deltas |
| Treat `phys_footprint` as the sole metric | Misses clean working-set latency and other processes | Worker, multi-process, and system view |
| Run `memory_pressure -S` during a benchmark | It simulates pressure rather than observing it | Plain `memory_pressure` snapshot and `vm_stat` deltas |
| Tune `madvise` for the pinned model | There is no mapped weight range | Delete unused allocations first |
| `mlock` the whole model | Increases non-reclaimable residency and can hit limits | Keep resident only what the product needs |
| Prune the GGML file before changing the loader | Current loader requires every expected tensor | Add explicit external-only load semantics first |
| Enable Core ML fallback after skipping encoder weights | Fallback has no weights to execute | Fail closed at initialization |
| Allocate a default state plus a state pool | Duplicates mutable runtime objects | `_no_state` initializer plus the minimum state count |
| Assume Q5 preserves quality | Quantization changes decoder arithmetic | Frozen exact-output and WER/pathology bakeoff |
| Add caller pools until SIGBUS disappears | Masks evidence and does not prove ownership | Crash, buffer-lifetime, and command-status diagnosis |
| Combine three compression mechanisms in one arm | Makes deltas and regressions unattributable | One changed mechanism per sealed arm |

## Works Cited

1. [Pinned whisper.cpp model loader and state implementation](https://github.com/ggml-org/whisper.cpp/blob/97c56f1dc1d1100a9d859c865a20c82d22f823ed/src/whisper.cpp) — custom GGML loading, backend weight allocation, external encoder selection, state caches, schedulers, and teardown.
2. [Pinned whisper.cpp public C API](https://github.com/ggml-org/whisper.cpp/blob/97c56f1dc1d1100a9d859c865a20c82d22f823ed/include/whisper.h) — context/state initializers, ownership, and same-context thread-safety boundary.
3. [Pinned whisper.cpp Core ML wrapper](https://github.com/ggml-org/whisper.cpp/blob/97c56f1dc1d1100a9d859c865a20c82d22f823ed/src/coreml/whisper-encoder.mm) — ARC requirement, `MLModel` lifetime, compute units, no-copy `MLMultiArray` input, and prediction pool.
4. [Pinned GGML Metal context](https://github.com/ggml-org/whisper.cpp/blob/97c56f1dc1d1100a9d859c865a20c82d22f823ed/ggml/src/ggml-metal/ggml-metal-context.m) — graph and transfer pools plus command-buffer ownership.
5. [Pinned GGML Metal device and buffer implementation](https://github.com/ggml-org/whisper.cpp/blob/97c56f1dc1d1100a9d859c865a20c82d22f823ed/ggml/src/ggml-metal/ggml-metal-device.m) — shared/private storage, no-copy wrapping, blits, and buffer teardown.
6. [Pinned whisper.cpp quantizer](https://github.com/ggml-org/whisper.cpp/blob/97c56f1dc1d1100a9d859c865a20c82d22f823ed/examples/quantize/quantize.cpp) — supported model layout and explicit skip list.
7. [Pinned common GGML quantization implementation](https://github.com/ggml-org/whisper.cpp/blob/97c56f1dc1d1100a9d859c865a20c82d22f823ed/examples/common-ggml.cpp) — supported types and two-dimensional-tensor rule.
8. [Official Whisper large-v3-turbo configuration](https://huggingface.co/openai/whisper-large-v3-turbo/blob/main/config.json) — architecture dimensions used in allocation arithmetic.
9. [Apple XNU `task_vm_info`](https://github.com/apple-oss-distributions/xnu/blob/main/osfmk/mach/task_info.h) — resident, compressed, footprint, peak, and neural ledger fields. The exact verified declarations also ship in the installed macOS 26.5 SDK.
10. [Apple Metal no-copy buffer API](https://developer.apple.com/documentation/metal/mtldevice/makebuffer(bytesnocopy:length:options:deallocator:)) — caller-provided storage contract.
11. Installed macOS 26.5 `footprint(1)`, `vmmap(1)`, `memory_pressure(1)`, `vm_stat(1)`, `mlock(2)`, and `madvise(2)` man pages — local primary platform contracts used to correct the raw report's accounting and VM advice.
