MissingManual

Whisper / memory · Collective Library

Low-Memory Whisper Turbo Inference on Apple Silicon

Collective Library edition. This is the complete technical report. Private filesystem paths, internal run identifiers, campaign-control notes, and repository navigation were removed. Technical claims, code, measurements, evidence labels, citations, corrections, and falsification criteria are preserved.

August 12, 2026

Executive Summary

The minimum-memory architecture is not “mmap the model and tune Darwin page hints.” The pinned whisper.cpp loader does not mmap the model at all. It opens the custom GGML .bin with std::ifstream, allocates backend buffers for every declared encoder and decoder tensor, and streams the weight bytes into those buffers. When Core ML is enabled, a separate MLModel performs encoder inference, but the native encoder weights remain allocated and loaded.

That establishes a deletion-first ladder:

  1. instrument the current resident FP16 hybrid and seal its real process and system footprint;
  2. run the stock whole-model Q5_0 control to learn the safe quantization and bandwidth effect;
  3. add an explicit fail-closed external-encoder load mode that never allocates or loads native encoder tensors;
  4. quantize only the remaining decoder-side weights after the lossless encoder-deletion control proves exact output parity;
  5. right-size one serialized state for the product's bounded audio and search contract; and
  6. compress the Core ML encoder only after the host-side waste is gone.

TASK_VM_INFO.phys_footprint is an important app-process metric, but it is not the sole truth and macOS has no documented universal “jetsam threshold” for a desktop process. A product receipt must pair the worker's footprint with RSS, dirty/clean/wired VM categories, relevant service processes, system compression and swap deltas, latency, and transcript parity.

The raw report's mandatory caller-side @autoreleasepool and implied SIGBUS cure are also unsupported for this pin. The pinned Metal backend already wraps graph execution and tensor-transfer paths in autorelease pools, retains and releases command buffers explicitly, and the Core ML wrapper uses ARC plus a prediction-local pool. Add another pool only when a measured application-level Objective-C loop accumulates autoreleased objects; do not treat it as a generic Metal crash fix.

1. Corrected Runtime Facts

1.1. Claims from the raw report that require correction

Raw-report claim Corrected verdict
Whisper weights are loaded through GGUF/GGML mmap False for the pinned runtime. Whisper uses its custom sequential GGML .bin, opens it with std::ifstream, allocates backend tensors, and reads bytes into them.
Clean mapped pages make model weights nearly free in phys_footprint Not applicable to the pinned loader. The weight buffers are allocated runtime storage, not a lazy read-only file mapping.
phys_footprint is the only metric that predicts termination Overstated. It is the task's charged footprint ledger and a primary metric for dirty memory impact. Clean working set, wired memory, opaque service allocations, compressor activity, swap, and system pressure still matter.
Core ML creates exactly two physical copies of the encoder Not established. The source proves a full native encoder allocation plus a separate Core ML model instance. Core ML's physical representation and accounting require measurement.
mlock converts clean file pages into charged footprint Imprecise and irrelevant here. mlock keeps mapped pages resident and is subject to wired and resource limits; the pinned model is not file-mapped.
MADV_WILLNEED is the recommended warm-model solution Not applicable here. madvise only advises an existing mapped range. The pinned model weights live in backend allocations.
MADV_FREE safely leaves model pages available for reuse Dangerous. MADV_FREE says the application no longer needs the contents and permits immediate reuse. It is not a model-weight cache hint.
Every inference loop needs a caller @autoreleasepool to avoid SIGBUS Unsupported. The pinned Metal and Core ML paths already contain scoped pools. SIGBUS has many causes; no primary source ties it to a missing outer pool here.
One context can serve parallel states with zero duplication Contradicted by the public header. The C interface says the same whisper_context must not be used concurrently. Every state also owns caches, schedulers, backends, and a Core ML model handle.
Setting audio_ctx=600 automatically shrinks cross-K/V allocation False for this pin. State initialization allocates cross and pad caches from the model's full n_audio_ctx, before request parameters set exp_n_audio_ctx.
Q5_0 has negligible accuracy loss and is the Apple Pareto frontier Unverified for this product. Upstream supports Q5_0; accuracy and latency remain model-, corpus-, and device-specific promotion gates.

1.2. The model file is not GGUF

The pinned Whisper loader expects its own ordered GGML binary. It checks GGML_FILE_MAGIC, then reads hyperparameters, mel filters, vocabulary, and named tensors. Vendored GGUF support elsewhere in the implementationsitory does not make this load path GGUF.

This distinction matters for memory engineering. Advice written for a GGUF mmap loader cannot be transferred to this Whisper runtime without inspecting the actual loader.

1.3. What Core ML replaces—and what it does not

Core ML replaces encoder compute after state initialization:

audio
  |
  v
mel preparation in whisper.cpp
  |
  v
Core ML encoder MLModel --------> encoder embeddings
                                      |
                                      v
                          whisper.cpp cross cache
                                      |
                                      v
                          whisper.cpp decoder

It does not currently replace native encoder storage. The model loader unconditionally constructs encoder tensor metadata, allocates all corresponding backend buffers, and loads every tensor before whisper_init_state opens the Core ML sidecar.

The source therefore proves structural redundancy: the native encoder tensor payload exists even though the native encoder graph is skipped. It does not prove that Core ML adds a byte-for-byte second FP16 physical copy. Core ML may compile, compress, cache, share, or charge its representation differently, and some allocations may reside in framework or service processes.

2. Allocation Topology

2.1. Model-global context

whisper_context owns:

  • model hyperparameters, vocabulary, and mel filters;
  • tensor metadata;
  • backend buffers containing all loaded model weights; and
  • optionally a default whisper_state when a non-_no_state initializer is used.

The pinned loader's decisive sequence is:

  1. create every encoder and decoder tensor description;
  2. allocate backend buffers for those descriptions;
  3. read every weight tensor directly into a host-capable buffer, or through a temporary buffer into another backend; and
  4. reject unknown, missing, or shape-mismatched tensors.

The file is closed after load. Deleting or advising file-cache pages cannot release the model's allocated backend storage.

2.2. Per-state allocations

Every whisper_state owns substantial mutable state:

  • backend instances;
  • self-attention, cross-attention, and flash-attention pad caches;
  • convolution, encoder, cross-attention, and decoder scheduler storage;
  • mel, logits, batch, candidate, and result vectors;
  • optional DTW and VAD structures; and
  • under WHISPER_USE_COREML, a separately initialized Core ML model handle.

When an external encoder is live, the pin omits the native encoder scheduler, but it still allocates the convolution shim, cross-attention scheduler, decoder scheduler, all caches, and all model weights.

Use whisper_init_from_file_with_params_no_state followed by exactly one whisper_init_state when custom state ownership is required. Calling the ordinary initializer and then allocating an additional state silently creates two states.

2.3. One serialized state is the memory floor

The public header's contract says the same context must not be used by multiple threads concurrently. whisper_full is explicitly not thread-safe for the same context. Separate states expose finer-grained APIs, but the header does not promote concurrent same-context execution as supported.

For a dictation product that finalizes one active speaker at a time, the minimum-memory topology is therefore simple:

one model context
        |
        +-- one state
                |
                +-- one serialized inference queue

Do not create a state pool until product concurrency requires it. Each state adds caches and schedulers and initializes another Core ML model object. Framework-internal sharing may reduce the incremental cost, but that is a measurement question, not a zero-duplication guarantee.

2.4. Metal shared buffers do not imply a CPU-plus-GPU copy

The pinned Metal backend can allocate CPU storage and wrap it with newBufferWithBytesNoCopy using shared storage. In that path the MTLBuffer references the same allocation; it is not evidence of a second full GPU copy. The backend keeps the underlying allocation alive with the Metal buffer and releases them together.

Other paths can use private Metal storage and explicit blits. Record the actual backend buffer names and storage path from the pinned build. “Unified memory” does not mean every API object aliases the same bytes, and “Metal buffer” does not by itself mean duplicate DRAM.

3. Exact State Arithmetic for Turbo

3.1. Architecture constants

The official openai/whisper-large-v3-turbo configuration and pinned GGML header use:

  • vocabulary: 51,866;
  • audio context: 1,500 positions;
  • model width: 1,280;
  • encoder layers: 32;
  • text context: 448 tokens; and
  • decoder layers: 4.

The model's FP16 file and loader-reported weight size are recorded in the local campaign note. File size is artifact evidence; it is not a peak-memory measurement.

3.2. K/V allocations use padded maxima

The pinned state initializer uses FP16 intermediate caches and rounds context lengths to multiples of 256. For Turbo:

Cache Source dimensions Allocated bytes
Self K and V 4 layers * 512 text slots * 1280 width * 2 arrays * 2 bytes 10,485,760
Cross K and V 4 layers * 1536 audio slots * 1280 width * 2 arrays * 2 bytes 31,457,280
Encoder pad K and V 1 layer * 1536 audio slots * 1280 width * 2 arrays * 2 bytes 7,864,320
Total All three caches 49,807,360

This corrects two common mistakes:

  1. use the padded 512 and 1,536 allocations, not logical 448 and 1,500; and
  2. include the separate pad cache.

Setting request audio_ctx=600 changes graph views after state creation. It does not change these allocations. A bounded-state patch would need to pass the product's maximum audio positions into state initialization and allocate GGML_PAD(max_audio_ctx, 256) deliberately.

That patch has a small ceiling compared with native encoder-weight deletion. At 600 positions, the rounded audio cache is 768 positions, so shrinking cross and pad caches can save only the difference between the 1,536- and 768-slot allocations. Prove it after the gigabyte-scale waste is removed.

3.3. Reserved address space is not physical footprint

The pin reserves n_vocab * n_text_ctx floats for the state's logits vector: 92,943,872 bytes for Turbo. std::vector::reserve obtains capacity, but Darwin does not necessarily charge every untouched demand-zero page to resident or physical footprint immediately. Runtime resize and tensor copies determine how much becomes touched.

Do not claim a 92.9 MB physical saving merely by removing the reserve. Sample before and after model load, state creation, first prefill, steady greedy decoding, and teardown. If the pages remain untouched, deleting the reserve may reduce virtual size without materially reducing peak inference RAM.

3.4. Compute buffers are empirical

The runtime logs scheduler sizes for convolution, encoder, cross-attention, and decoder paths. Those values depend on backend, flash attention, graph shape, and build. Capture the log lines in the receipt instead of deriving them from parameter count.

An external encoder should remove the native encoder scheduler. It does not remove:

  • the mel/conv shim;
  • encoder output storage;
  • cross-attention K/V construction;
  • decoder compute storage; or
  • native encoder weights in the current loader.

4. Darwin Memory Accounting

4.1. No single column is enough

The installed footprint(1) documentation defines a process footprint around dirty memory charged to that process and recommends reducing dirty memory first. It separately reports clean, reclaimable, swapped, and wired memory. Clean file-backed working sets can still cause latency thrash; wired memory cannot be paged, compressed, or reclaimed.

Use a metric set rather than a winner:

Metric What it answers Important blind spot
TASK_VM_INFO.phys_footprint Current footprint charged to the worker task Does not attribute opaque work in other processes or describe clean working-set latency
ledger_phys_footprint_peak Peak charged footprint observed for the worker Requires a sufficiently new TASK_VM_INFO revision; still worker-only
resident_size / RSS Pages currently resident for the process Includes categories with different reclaimability and can overstate durable pressure
footprint dirty/clean/reclaimable/wired columns Category-level task or multi-process accounting Cross-process inspection can require root; sampling has overhead
vmmap -summary VM regions, dirty/swapped allocation, and mapped-file categories Snapshot, not a peak; permissions may restrict inspection
vm_stat deltas System compressor, page-in/out, and swap activity System-wide and sensitive to unrelated work
memory_pressure snapshot Current system memory statistics and free percentage Not per-process; -S simulates pressure and invalidates a benchmark
Core ML/ANE service census Whether work appears outside the worker Process names and accounting can change by OS version; causality needs a control

macOS desktop process termination is not specified as one public footprint threshold. A low worker footprint can coexist with high system pressure, and a high RSS can include reclaimable or clean pages. Report both.

4.2. In-process sampler

The installed SDK exposes current and peak fields through TASK_VM_INFO. Sample in-process so the worker can bind memory points to exact inference events:

#include <mach/mach.h>
#include <mach/task_info.h>

#include <cstdint>

struct DarwinTaskMemory {
    uint64_t resident_bytes;
    uint64_t resident_peak_bytes;
    uint64_t compressed_bytes;
    uint64_t footprint_bytes;
    uint64_t footprint_peak_bytes;
};

static bool read_darwin_task_memory(DarwinTaskMemory & out) {
    task_vm_info_data_t info = {};
    mach_msg_type_number_t count = TASK_VM_INFO_COUNT;
    const kern_return_t result = task_info(
        mach_task_self(),
        TASK_VM_INFO,
        reinterpret_cast<task_info_t>(&info),
        &count);

    if (result != KERN_SUCCESS || count < TASK_VM_INFO_REV1_COUNT) {
        return false;
    }

    out.resident_bytes = info.resident_size;
    out.resident_peak_bytes = info.resident_size_peak;
    out.compressed_bytes = info.compressed;
    out.footprint_bytes = info.phys_footprint;
    out.footprint_peak_bytes =
        count >= TASK_VM_INFO_REV3_COUNT &&
                info.ledger_phys_footprint_peak > 0
            ? static_cast<uint64_t>(info.ledger_phys_footprint_peak)
            : 0;
    return true;
}

Treat a zero peak as unavailable unless the surrounding API call failed. Do not silently replace it with the current value.

4.3. Required measurement phases

For every candidate arm, capture the same event sequence:

  1. quiet-host admission and system baseline;
  2. worker process started, before model load;
  3. after native model context load;
  4. after the single state and Core ML model load;
  5. after first specialization or warmup;
  6. immediately before each measured request;
  7. at a high enough cadence during the request to observe the peak;
  8. immediately after final text;
  9. after an idle interval; and
  10. after explicit state/context teardown.

A before/after snapshot is insufficient when an intermediate compute buffer is the peak. Prefer in-process phase markers plus a 50-100 ms external sampler. Record sampler overhead with a no-sampler latency control.

4.4. Process and system scope

At minimum, preserve:

# Per-process VM category snapshot.
vmmap -summary "$WORKER_PID"

# Current RSS in KiB; not a peak and not a footprint substitute.
ps -o pid=,rss=,vsz=,command= -p "$WORKER_PID"

# System totals followed by one-second deltas.
vm_stat 1

# Current system memory statistics. Do not add -S to a measurement run.
memory_pressure

Use footprint --sample when privileges and overhead are acceptable. When comparing multiple related processes, use its multi-process de-duplication rather than summing RSS columns blindly.

Core ML may perform work in framework or service processes that are not child processes of the app. Record a process census before and during the Core ML arm and pair it with a native-encoder control. Do not hard-code a service name as a stable API or attribute every system-service delta to the model.

5. The Deletion-First Optimization Ladder

5.1. Stage 0: seal the existing FP16 hybrid

Before changing allocation, measure the exact current product candidate. The receipt must bind:

  • model file SHA-256 and byte size;
  • Core ML package and compiled-tree identity;
  • whisper.cpp commit and dirty-source state;
  • build flags, linked libraries, and compute-unit policy;
  • process and system memory series;
  • runtime allocation log lines;
  • p50, p95, and maximum stop-to-final latency; and
  • exact tokens, EOT, final text, and quality-slice results.

This stage converts “approximately 2-3 GB” into a physical baseline. Without it, later deltas cannot be attributed.

5.2. Stage 1: stock whole-model Q5_0

The pinned quantizer supports Q5_0 and applies it to eligible two-dimensional tensors. It deliberately skips encoder convolution biases, encoder positional embeddings, and decoder positional embeddings; non-two-dimensional tensors are also retained.

Q5_0 is a high-leverage first control because the loader stores weights in their quantized tensor type. It can reduce the large native backend buffer and memory bandwidth without a custom runtime patch.

It is not lossless. Require:

  • artifact/header validation and exact source binding;
  • model-load success with the same Core ML sidecar;
  • latency and memory measured with the same protocol;
  • exact-output diagnostics on frozen rows; and
  • full WER and product-pathology qualification before promotion.

Do not state an expected byte size or accuracy delta before the generated artifact and bakeoff exist.

5.3. Stage 2: skip native encoder allocation under fail-closed Core ML

This is the largest source-verified deletion opportunity. The lossless control should retain:

  • the original header, vocabulary, filters, and decoder tensors;
  • decoder embeddings and output projection;
  • all decoder self- and cross-attention weights;
  • the external encoder output ABI; and
  • a hard requirement that the exact Core ML model loads successfully.

It should omit backend allocation for:

  • encoder positional embeddings;
  • both encoder convolution layers; and
  • all 32 native encoder transformer layers.

Do not begin by inventing a new “decoder-only GGML” container. The smaller first patch can keep the existing full artifact and teach an explicit external-encoder-only loader mode to validate encoder tensor names and shapes while seeking or discarding their bytes without allocating backend tensors. That isolates RAM deletion from file-format surgery.

The current generic loader callback has read, eof, and close, but no seek/skip callback. Reading a skipped giant tensor into a temporary vector would recreate the peak being removed. Add either:

  1. a bounded chunk discard path; or
  2. an optional checked skip callback implemented by the file loader.

The skip path must reject truncated input, overflow, unexpected tensor names, wrong shapes, wrong types, missing tensors, duplicate tensors, and a Core ML load failure. It must never silently fall back to native encoding because the native weights do not exist.

The external encoder call must fail closed too. A successful model load is not enough: propagate predictionFromFeatures errors, require exactly one Float32 [1, 600, 1280] output, validate its element count and strides before copying, and return an encoding failure rather than decoding stale or malformed state. An unchecked memcpy(output.count) is incompatible with a runtime whose native encoder fallback has deliberately been deleted.

Promotion gate: exact token IDs, EOT, text, logits within a preregistered tolerance if captured, and latency parity against the full FP16 hybrid. The memory delta is then causally attributable to unused native encoder storage.

5.4. Stage 3: quantize the retained decoder

Only after Stage 2 proves lossless deletion should quantization be applied to the remaining eligible decoder tensors. This cleanly separates:

  • exact-output changes caused by deleting unused tensors—which should be none; from
  • numerical changes caused by decoder quantization.

A compact decoder-only shipping artifact can follow later if disk size and load I/O justify a new format. The runtime patch and its validation rules are the hard part; truncating bytes first merely creates an incompatible file.

5.5. Stage 4: right-size mutable state

After weight waste is gone, test smaller allocations in descending value:

  1. exactly one state and one serialized request;
  2. greedy search with one decoder rather than beam candidates;
  3. DTW/token-timestamp structures disabled when the product does not consume them;
  4. bounded cross and pad caches created from the product's maximum external encoder positions; and
  5. logits capacity and decoder work buffers sized from observed maximum batch rather than theoretical text context.

Every bound needs a fail-closed input check. A 600-position state must reject an encoder output longer than 600 rather than overrun, silently truncate, or reuse stale cache tails.

5.6. Stage 5: compress Core ML last

Core ML palettization can reduce package storage and may reduce runtime weight cost, but package bytes are not resident bytes. The compiler can transform or expand representations, and compute placement can change.

For each compression arm, separately prove:

  • portable package size and compiled-tree size;
  • numerical parity against the dense encoder;
  • end-to-end transcript and WER behavior;
  • actual ANE placement;
  • cold specialization and warm-load behavior;
  • app and service-process memory; and
  • stop-to-final latency under the same contention contract.

Do not combine decoder quantization, native-encoder deletion, and Core ML compression in the first experiment. One changed mechanism per arm preserves causal evidence.

5.7. Idle residency is a product policy, not a kernel trick

Loading while the user speaks can hide resident initialization from stop-to-final latency. Releasing the context after an idle timeout can lower idle RAM. Neither reduces peak inference memory.

Measure three policies:

Policy User benefit Memory cost Risk
Always resident Lowest request-start variance Highest idle RAM System pressure while unused
Load on recording start Hides load behind speech Peak unchanged, idle lower Short utterances may stop before load finishes
Idle timeout Amortizes nearby requests Tunable First request after timeout pays load/specialization

Core ML's first device specialization is a separate product event. Do not confuse install-time compilation, warm model load, and steady prediction.

6. mmap, mlock, and madvise

6.1. Why they are not current levers

The pinned Whisper path does not expose a mapped model-weight range. Therefore:

  • there is no model mapping to lock;
  • there is no model mapping to prefetch with MADV_WILLNEED;
  • dropping the source file's cache does not free backend weights; and
  • a benchmark of file-cache advice would test I/O, not resident inference RAM.

Delete this work from the immediate plan.

6.2. If a future loader introduces mmap

Treat these APIs according to their platform contracts:

  • mlock keeps the addressed pages resident until unlocked and is bounded by process and system limits. It increases non-reclaimable pressure and should be an explicit latency-versus-memory experiment, not a default.
  • MADV_WILLNEED is advisory. The man page does not promise asynchronous completion, full prefetch, or future residency.
  • MADV_DONTNEED says the range is not expected soon; it is not a durability or latency guarantee.
  • MADV_FREE permits the VM to reuse contents. Never apply it to live model weights.

An mmap rewrite might reduce dirty weight storage by relying on clean file-backed pages, but Metal compatibility, tensor alignment, quantized access, page-fault tails, and Core ML coexistence all need a separate design and physical falsifier. It is not prerequisite to removing unused encoder tensors.

7. Autorelease Pools, No-Copy Buffers, and SIGBUS

7.1. What the pin already does

The pinned sources establish:

  • src/coreml/whisper-encoder.mm refuses to compile without ARC;
  • Core ML prediction is wrapped in @autoreleasepool;
  • Metal graph compute and tensor transfer paths contain their own @autoreleasepool blocks;
  • command buffers created with unretained references are explicitly retained, replaced, synchronized, and released by the backend; and
  • shared Metal buffers can wrap backend-owned CPU storage with a nil deallocator while that storage remains owned by the backend buffer.

That is not evidence of a pool leak.

7.2. What an outer pool can and cannot do

An application-level @autoreleasepool around a long Objective-C loop can bound objects created by the application or frameworks and released by autorelease. Use it when an allocation trace shows growth between pool drains.

It cannot repair:

  • a CPU pointer freed before an asynchronous GPU operation completes;
  • out-of-bounds tensor views or incorrect strides;
  • a command buffer or resource that is explicitly retained and never released;
  • invalid model bytes;
  • an OS or driver fault; or
  • memory pressure caused by live, required buffers.

7.3. No-copy lifetime rule

Apple's no-copy Metal initializer wraps caller-provided bytes. The caller must keep those bytes valid for the lifetime and use of the MTLBuffer. In an asynchronous path, lifetime must extend through GPU completion, not merely the encoding call.

When SIGBUS occurs, preserve the crash report, fault address, GPU error state, command-buffer status, tensor name and range, buffer ownership, and completion ordering. Do not diagnose “missing autorelease pool” from the signal alone.

8. Decision Framework

Mechanism Expected memory target Accuracy risk Latency risk Evidence status
Measure current FP16 hybrid None; establishes truth None Sampling overhead Required first
Whole-model Q5_0 Native backend weights Real; must qualify May improve bandwidth or add dequant cost Source-supported candidate
Skip native encoder allocation Unused native encoder weights None if external path is exact and fail-closed Load path changes; inference should match Highest-value hypothesis from source
Quantize retained decoder Decoder weights and bandwidth Real; must qualify Device-dependent Test after lossless deletion
Single _no_state context plus one state Duplicate caches, schedulers, Core ML objects None for serialized product Removes concurrency Source-verified topology
Bound cross/pad cache to 600 Tens of MB, not GB Input-bound failure if contract violated Small possible gain Secondary source-supported patch
Remove oversized logits reserve Virtual capacity; physical saving uncertain Possible batch regression Probably small Measure before implementation
Disable DTW and beam search Optional mutable state Product-feature and decode-quality tradeoff Usually favorable Product-contract dependent
LUT-compress Core ML encoder Core ML package and possibly runtime weights Numerical and placement risk Compile and prediction risk Last-stage physical experiment
Idle unload Idle memory only None Reload/specialization tail Product policy
Add model mmap Dirty weight storage, possibly Loader/backend rewrite risk Page-fault tails Separate future program
Add mlock/madvise now None in current loader None Engineering distraction Delete

9. Failure Modes and Debugging Playbook

9.1. File size falls but peak footprint does not

Likely causes:

  • the quantized file still expands into another backend type;
  • Core ML or compute buffers dominate the peak;
  • the sampler missed the true peak;
  • both old and new contexts overlap during model replacement; or
  • allocator high-water pages remain charged after a transient conversion.

Check backend buffer log sizes, phase markers, context overlap, and footprint/vmmap categories before changing quantization again.

9.2. Worker footprint falls but system pressure does not

Run the native-encoder and Core ML controls with identical host admission. Inspect relevant service-process and system deltas. Core ML work may be charged outside the worker, but unrelated services can move at the same time. Require repeatable paired deltas before attribution.

9.3. A second state causes a large jump

Verify the initializer. If the ordinary context initializer created a default state and the application then called whisper_init_state, free the duplicate and switch to the _no_state initializer. Next compare state creation with and without Core ML to separate caches/schedulers from framework model residency.

9.4. audio_ctx=600 did not shrink initialization logs

This is expected for the pin. The request parameter is assigned after state creation, while K/V caches were allocated from full model hyperparameters. Implement a checked state-init bound or stop claiming a memory saving.

9.5. RSS is much larger than footprint

Use footprint and vmmap to split dirty, clean, reclaimable, wired, and shared categories. Do not optimize the numerical difference itself. Optimize the category that limits the product: charged pressure, clean working-set thrash, wired allocations, or latency after eviction.

9.6. Memory grows across repeated requests

Separate three patterns:

  1. one-time lazy allocation that plateaus;
  2. allocator caching that remains reusable; and
  3. unbounded growth proportional to request count.

Run enough iterations to distinguish them, add explicit pool-drain markers at the application boundary, and record live object/VM categories. A stable higher plateau is not automatically a leak; linear growth is not automatically an autorelease problem.

9.7. Q5 changes one proper noun or number

That is an accuracy failure for exact-output parity even if aggregate WER barely moves. Keep the full pathology matrix: names, numbers, negations, punctuation, false starts, repetitions, silence, and clean dictation. Promote using the product's quality policy, not file size or average WER alone.

9.8. Metal faults after a no-copy change

Audit ownership and ordering first:

  1. identify the exact pointer and byte range;
  2. prove alignment and buffer bounds;
  3. prove the CPU allocation remains alive;
  4. wait for or otherwise order GPU completion before reuse/free;
  5. inspect command-buffer error details; and
  6. rerun with a copied-buffer control.

An outer autorelease pool is not a substitute for those proofs.

10. Production Qualification Receipt

Every memory candidate should emit one sealed record with:

Artifact identity

  • source model and revision;
  • GGML model SHA-256, byte size, and ftype;
  • Core ML package and compiled-tree digests;
  • whisper.cpp commit, patch digest, and dirty status;
  • executable and linked-runtime digests; and
  • OS, hardware model, RAM capacity, and boot identity.

Runtime topology

  • context initializer used;
  • state count;
  • Core ML model count;
  • backend buffer names and logged sizes;
  • compute-unit setting;
  • flash-attention, timestamp, VAD, search, and audio-context settings; and
  • serialized or concurrent request policy.

Memory series

  • worker current and peak phys_footprint;
  • worker current and peak resident size;
  • worker compressed bytes;
  • footprint category summaries when available;
  • relevant service-process series or an explicit “not attributable” marker;
  • system wired/compressor/page-in/page-out/swap deltas; and
  • baseline, post-load, post-state, warmup, per-request peak, idle, and teardown phase timestamps.

Product outcomes

  • stop-to-final p50, p95, and maximum;
  • cold load and first-specialization time reported separately;
  • exact token IDs, EOT, and final text for diagnostic fixtures;
  • WER and pathology gates on frozen slices;
  • silence false alarms;
  • memory growth slope over a long repeated-request soak; and
  • cancellation, fallback, and idle-unload behavior.

The promotion rule is conjunctive: lower memory is useful only if latency, quality, stability, and fail-closed behavior still pass.

11. Anti-Patterns

Anti-pattern Why it fails Use instead
Quote model file size as RAM Ignores allocator, runtime, state, services, and compression Phase-aligned physical measurements
Treat RSS as the sole metric Mixes categories with different reclaimability Footprint plus VM categories and system deltas
Treat phys_footprint as the sole metric Misses clean working-set latency and other processes Worker, multi-process, and system view
Run memory_pressure -S during a benchmark It simulates pressure rather than observing it Plain memory_pressure snapshot and vm_stat deltas
Tune madvise for the pinned model There is no mapped weight range Delete unused allocations first
mlock the whole model Increases non-reclaimable residency and can hit limits Keep resident only what the product needs
Prune the GGML file before changing the loader Current loader requires every expected tensor Add explicit external-only load semantics first
Enable Core ML fallback after skipping encoder weights Fallback has no weights to execute Fail closed at initialization
Allocate a default state plus a state pool Duplicates mutable runtime objects _no_state initializer plus the minimum state count
Assume Q5 preserves quality Quantization changes decoder arithmetic Frozen exact-output and WER/pathology bakeoff
Add caller pools until SIGBUS disappears Masks evidence and does not prove ownership Crash, buffer-lifetime, and command-status diagnosis
Combine three compression mechanisms in one arm Makes deltas and regressions unattributable One changed mechanism per sealed arm

Works Cited

  1. Pinned whisper.cpp model loader and state implementation — custom GGML loading, backend weight allocation, external encoder selection, state caches, schedulers, and teardown.
  2. Pinned whisper.cpp public C API — context/state initializers, ownership, and same-context thread-safety boundary.
  3. Pinned whisper.cpp Core ML wrapper — ARC requirement, MLModel lifetime, compute units, no-copy MLMultiArray input, and prediction pool.
  4. Pinned GGML Metal context — graph and transfer pools plus command-buffer ownership.
  5. Pinned GGML Metal device and buffer implementation — shared/private storage, no-copy wrapping, blits, and buffer teardown.
  6. Pinned whisper.cpp quantizer — supported model layout and explicit skip list.
  7. Pinned common GGML quantization implementation — supported types and two-dimensional-tensor rule.
  8. Official Whisper large-v3-turbo configuration — architecture dimensions used in allocation arithmetic.
  9. Apple XNU task_vm_info — resident, compressed, footprint, peak, and neural ledger fields. The exact verified declarations also ship in the installed macOS 26.5 SDK.
  10. Apple Metal no-copy buffer API — caller-provided storage contract.
  11. Installed macOS 26.5 footprint(1), vmmap(1), memory_pressure(1), vm_stat(1), mlock(2), and madvise(2) man pages — local primary platform contracts used to correct the raw report's accounting and VM advice.