MissingManual

Core ML / allocation · Collective Library

Core ML Preallocated Multi-Output Buffers on Apple Silicon

Collective Library edition. This is the complete technical report. Private filesystem paths, internal run identifiers, campaign-control notes, and repository navigation were removed. Technical claims, code, measurements, evidence labels, citations, corrections, and falsification criteria are preserved.

August 9, 2026

Executive Summary

MLPredictionOptions.outputBackings is a public way to propose destination objects for Core ML outputs. It is not a direct-write command. Core ML may ignore a proposed backing when the model does not support it or when prediction uses batch mode, and an unknown output name is silently ignored. A successful prediction is therefore insufficient evidence. The caller must prove that the returned output is the exact proposed object.

For a multi-array output, the backing must be an MLMultiArray accepted by the output's MLFeatureDescription.isAllowedValue(_:). Apple explicitly recommends a page-aligned address for a raw external-memory MLMultiArray. The public initializer wraps that storage without copying when the wrapper is created. Neither fact proves that the selected compute device writes directly into those bytes; Core ML may still stage internally.

An IOSurface-backed CVPixelBuffer, wrapped as an MLMultiArray, can reduce inference latency by avoiding a buffer copy. That is the SDK's deliberately conditional wording. It is not a blanket zero-copy guarantee. There is no public MLFeatureValue initializer over MTLBuffer in the installed SDK.

For one Whisper family invocation, the two outputs are K and V with the same shape. Turbo uses two [4, 1536, 1280] FP16 outputs; Distil uses two [2, 1536, 1280] FP16 outputs. These are alternative packages. A single invocation never combines one Turbo-shaped output with one Distil-shaped output.

Treat direct output ownership as a falsifiable optimization:

  1. validate names, feature types, shapes, data types, and strides;
  2. run one synchronous prediction with distinct K and V backings;
  3. require returned-object identity for both outputs;
  4. prove numerical parity, complete writes, guard preservation, and zero tails;
  5. measure allocation, packing, copying, and end-to-end warmed latency; and
  6. keep placement evidence separate from backing acceptance.

1. The Contract: Proposal, Validation, and Proof

1.1. Public API and version matrix

API or behavior Availability in installed SDK What is documented What is not documented
MLPredictionOptions.outputBackings macOS 11, iOS 16, watchOS 9, tvOS 16 Dictionary from feature name to MLMultiArray or CVPixelBuffer, according to feature type Acceptance by every mlprogram, direct device writes, or a latency saving
MLMultiArray.init(dataPointer:shape:dataType:strides:deallocator:) Public MLMultiArray API Wraps existing storage without a copy at construction No internal staging during prediction or a documented deallocator queue
MLMultiArray.init(pixelBuffer:shape:) macOS 12, iOS 16, watchOS 9, tvOS 16 Wraps and owns an IOSurface-backed pixel buffer; can reduce latency by avoiding a copy Zero copies for every model, compute unit, OS, shape, or prediction mode
MLFeatureDescription.isAllowedValue(_:) Public Core ML API Validates a feature value against its declared output description That passing validation forces Core ML to use the backing
Batch prediction Public Core ML API outputBackings may be ignored in batch mode Stable object reuse or caller-controlled output allocation for a batch
MLFeatureValue over MTLBuffer No public initializer in installed SDK Nothing Direct public Metal-buffer binding to a model output

The outputBackings header defines five fail-closed rules:

  • a missing feature entry permits framework allocation;
  • an unsupported model permits framework allocation;
  • batch prediction permits framework allocation;
  • an unknown feature name is ignored; and
  • a known backing must satisfy the output feature description's isAllowedValue(_:) test or prediction reports an error.

After prediction, compare the returned multi-array to the proposed backing by object identity. Equal shapes, equal pointers observed at different times, or equal bytes are weaker claims. They do not prove that Core ML accepted the proposed object.

1.2. What “without copy” actually establishes

The raw-pointer initializer's “without copy” statement is about constructing the MLMultiArray wrapper around existing storage. It establishes that the wrapper does not clone the initial bytes. It does not specify how CPU, GPU, or Neural Engine execution reaches that storage.

Likewise, the pixel-buffer initializer says an IOSurface-backed array can reduce inference latency by avoiding a buffer copy. It does not promise that every output is written in place, that no temporary exists, or that an accepted backing preserves Neural Engine placement.

The categorical claim that the Neural Engine always copies to plain malloc storage is therefore unsupported. It also conflicts with Apple's explicit recommendation to use a page-aligned raw MLMultiArray address for best outputBackings performance. Test both mechanisms on the target model.

2. Family-Specific Fixed Shapes

The shape contract must be derived from one model package, not assembled across families.

Family package K shape V shape C-order element strides Bytes per output K plus V
Turbo [4, 1536, 1280] [4, 1536, 1280] [1966080, 1280, 1] 15,728,640 30 MiB
Distil [2, 1536, 1280] [2, 1536, 1280] [1966080, 1280, 1] 7,864,320 15 MiB

The first stride is independent of layer count because one layer spans 1536 * 1280 = 1,966,080 elements. FP16 uses two bytes per element. The runtime contract must compare the actual output description and returned array against these values; arithmetic in a guide is not an ABI.

The deployed cross-attention cache uses only positions 0:1500. Positions 1500:1536 are physical padding and must remain bitwise FP16 zero in the campaign candidate. A zeroed destination before prediction is not sufficient: the post-prediction verifier must inspect every tail byte for K and V.

2.1. Pixel-buffer layout

For MLMultiArray.init(pixelBuffer:shape:), the installed header documents an exact mapping: the last shape dimension equals pixel-buffer width, and the product of all preceding dimensions equals height. Therefore a Turbo backing would use width 1280 and height 6144; a Distil backing would use width 1280 and height 3072.

The pixel format must be kCVPixelFormatType_OneComponent16Half for FP16. A consumer must still respect CVPixelBufferGetBytesPerRow; Core Video may pad rows. Do not turn the documented shape mapping into an assumption of dense row bytes.

3. Buffer Mechanisms

Mechanism Documented fact Working interpretation
Page-aligned raw storage plus MLMultiArray Apple recommends page alignment; initializer wraps existing storage without copying Simplest C/C++ ownership path and the first mechanism to test
IOSurface-backed CVPixelBuffer wrapped as MLMultiArray Can reduce latency by avoiding a copy; wrapper owns the pixel buffer Strong alternative when raw storage is ignored or slower, but still requires identity and timing proof
CVPixelBufferPool Apple calls an IOSurface-backed pool often efficient, especially for playback or export May reduce repeated allocation overhead; relevance to large tensor outputs is unverified
Framework-allocated MLMultiArray Always available as the control Safest ownership baseline, with an explicit checked copy into runtime cache
MTLBuffer as MLFeatureValue No public initializer exists in installed MLFeatureValue.h Do not ship a private or invented bridge
Accelerate over returned bytes Accelerate can process caller-accessible memory after prediction Useful post-processing, but not evidence that prediction was zero-copy

IOSurface can reduce copies across compatible framework boundaries. It does not make “zero-copy” an API property. Promote it only after the exact target model, OS, compute units, and prediction path prove identity, parity, and a measurable end-to-end saving.

4. Ownership, Lifetime, and Concurrency

Object Caller obligation Public guarantee boundary
Raw allocation Keep bytes valid until the wrapper is deallocated Deallocator is called when the array is deallocated; its thread is unspecified
MLMultiArray backing Keep the object alive through prediction and result validation Core ML may return a different array
Pixel buffer Do not lock its base address during prediction Pixel-buffer-backed array owns and releases the buffer
K and V backings Use distinct storage and validate both returned identities No public output-write ordering or aliasing guarantee
Reused backing Serialize access until the prior synchronous prediction and reads finish No public guarantee for overlapping predictions into one object

The installed headers do not promise a queue or thread for the external-memory deallocator. Write the closure so it has no main-thread requirement, captures only stable ownership, and safely releases exactly once. Do not state that Core ML uses one background deallocation thread.

The public API also does not specify whether multiple outputs are produced serially or concurrently. Never alias K and V, never alias an input and output, and do not overlap predictions that reuse the same storage. These rules are conservative caller policy, not claims about Core ML's hidden scheduler.

An @autoreleasepool around a long-running Objective-C loop can bound temporary Objective-C objects. Community reports associate unbounded loops with IOSurface pressure, but the installed headers do not make one autoreleasepool per prediction mandatory or guarantee that it cures every pool exhaustion failure. Treat it as an operational experiment with resident-memory telemetry.

5. Minimal Implementation Patterns

5.1. Objective-C++ page-aligned raw backings

This sketch uses the SDK-recommended raw-memory path. It intentionally does not call it a direct Neural Engine write.

#import <CoreML/CoreML.h>
#import <mach/vm_page_size.h>

static MLMultiArray *MakeFP16Backing(NSArray<NSNumber *> *shape,
                                      size_t elementCount,
                                      NSError **error) {
    void *bytes = NULL;
    const size_t byteCount = elementCount * sizeof(uint16_t);
    if (posix_memalign(&bytes, vm_page_size, byteCount) != 0) {
        return nil;
    }

    MLMultiArray *array = [[MLMultiArray alloc]
        initWithDataPointer:bytes
                      shape:shape
                   dataType:MLMultiArrayDataTypeFloat16
                    strides:@[@1966080, @1280, @1]
                deallocator:^(void *ownedBytes) { free(ownedBytes); }
                      error:error];
    if (array == nil) {
        free(bytes);
    }
    return array;
}

NSArray<NSNumber *> *shape = @[@4, @1536, @1280]; // Turbo only.
const size_t elementCount = 4ULL * 1536ULL * 1280ULL;
MLMultiArray *crossK = MakeFP16Backing(shape, elementCount, &error);
MLMultiArray *crossV = MakeFP16Backing(shape, elementCount, &error);

MLFeatureDescription *kDescription =
    model.modelDescription.outputDescriptionsByName[@"cross_k"];
MLFeatureDescription *vDescription =
    model.modelDescription.outputDescriptionsByName[@"cross_v"];

BOOL kAllowed = [kDescription
    isAllowedValue:[MLFeatureValue featureValueWithMultiArray:crossK]];
BOOL vAllowed = [vDescription
    isAllowedValue:[MLFeatureValue featureValueWithMultiArray:crossV]];
if (!kAllowed || !vAllowed) {
    // Reject the candidate before prediction; do not infer acceptance.
}

MLPredictionOptions *options = [[MLPredictionOptions alloc] init];
options.outputBackings = @{@"cross_k": crossK, @"cross_v": crossV};
id<MLFeatureProvider> result =
    [model predictionFromFeatures:input options:options error:&error];

MLMultiArray *returnedK =
    [result featureValueForName:@"cross_k"].multiArrayValue;
MLMultiArray *returnedV =
    [result featureValueForName:@"cross_v"].multiArrayValue;
if (returnedK != crossK || returnedV != crossV) {
    // Core ML did not accept both proposals. Use the checked-copy control.
}

Production code should centralize the family shape and byte count, check every allocation and feature lookup, retain the backings for the full use interval, and make failure select an explicit control path. Never continue as if an ignored backing were accepted.

5.2. Swift identity check

The same proof is required from Swift:

let options = MLPredictionOptions()
options.outputBackings = ["cross_k": crossK, "cross_v": crossV]

let result = try model.prediction(from: input, options: options)
guard result.featureValue(for: "cross_k")?.multiArrayValue === crossK,
      result.featureValue(for: "cross_v")?.multiArrayValue === crossV else {
    throw BackingError.outputBackingIgnored
}

Do not use batch prediction for the first acceptance proof. The public header explicitly names batch as a mode where the proposed backing may not be used.

5.3. IOSurface-backed alternative

When testing the pixel-buffer path:

  1. create an IOSurface-backed CVPixelBuffer with the exact width, height, and kCVPixelFormatType_OneComponent16Half format;
  2. wrap it with MLMultiArray.init(pixelBuffer:shape:);
  3. do not lock the base address or call array data/subscript accessors before prediction;
  4. propose the wrapper through outputBackings;
  5. require returned MLMultiArray identity; and
  6. lock and inspect bytes only after synchronous prediction completes.

The wrapper owns the pixel buffer. Audit retains and releases accordingly; do not release storage as though ownership remained solely with the creator.

6. Falsification Recipe

Separate five questions that are easy to blur together:

Question Required evidence
Is the API available? Compile against the target SDK and record deployment target
Is this value admissible? Exact output name plus isAllowedValue(_:) success
Was this object used? Returned K and V object identity
Is the result correct? Shape, type, stride, range, finite-value, parity, guard, digest, and tail checks
Did the product get faster? Paired warmed stage and WAV-to-final-text timings under a quiet-host gate

6.1. Acceptance and corruption checks

For every scored prediction:

  • require the exact model output-name set before constructing options;
  • require the expected feature type, shape, data type, and element strides;
  • require isAllowedValue(_:) for each proposed backing;
  • require returned-object identity for both K and V;
  • require the returned pointer range to stay within the owned allocation;
  • compare all logical elements against the framework-allocation control under a frozen numerical tolerance;
  • hash all logical output bytes, not a prefix sample;
  • place canaries around caller-owned storage and require them unchanged;
  • require every padded tail byte to remain zero; and
  • reject NaN, infinity, all-zero logical tensors, stale prior-call digests, and K/V aliasing.

Object identity proves that Core ML returned the proposed object. It does not prove there was no hidden staging before the result arrived. Latency and allocation evidence address that separate hypothesis.

6.2. Timing and allocation A/B

Compare exactly three arms with the same compiled model and inputs:

  1. framework output plus explicit checked copy;
  2. page-aligned raw MLMultiArray output backing; and
  3. IOSurface-backed MLMultiArray output backing.

Keep model load, warmup count, compute units, thread policy, input order, and host state fixed. Report prediction, packing/copy, complete encoder stage, and WAV-to-final-text time separately. Include allocation and wrapper creation in the interval unless the product reuses those objects persistently.

Use allocation tracing, Time Profiler, signposts around the owned stages, and a compute-plan inspection as supporting evidence. The public SDK does not guarantee an Instruments lane that exposes every Core ML copy, nor does absence of a visible copy event prove zero-copy execution. Static compute-unit choice also does not prove Neural Engine residency or specify fallback behavior.

6.3. Decision rules

  • Reject API feasibility if either output name is absent, either proposed value is disallowed, or either returned identity differs.
  • Reject correctness on any parity, range, stride, canary, tail, stale-data, or termination failure.
  • Reject copy-removal language unless the accepted backing produces a repeatable paired stage saving beyond timer noise and allocation effects.
  • Reject product promotion unless the saving survives the campaign's full end-to-end latency and quality gates.
  • Retain the explicit-copy control whenever the runtime, OS, compiled model, family, or output description changes.

7. Failure Modes

Symptom Likely cause Fail-closed response
Prediction succeeds but identity differs Unsupported backing, batch mode, or omitted entry Score as ignored; use checked copy
Misspelled output appears harmless Unknown names are ignored Compare exact output-name sets before prediction
Prediction reports an admissibility error Backing fails feature description Reject shape, type, or storage choice; do not coerce silently
Pixel-buffer path stalls or fails Base address or array data access locked it before prediction Recreate unlocked backing and enforce access order
Correct shape but corrupted cache Wrong element strides or ignored row bytes Validate strides and use CVPixelBufferGetBytesPerRow
Tail contains nonzero values Model/package wrote padding or verifier inspected wrong range Reject artifact; never mask after the fact
Repeated calls return stale bytes Overlap, lifetime error, or partial write Serialize, seed distinct canaries/digests, and reject repeats
Candidate allocates less but is not faster Allocation was not the bottleneck or staging remains Keep control and stop optimization
Placement changes with backing Runtime selected a different execution plan Treat as a separate placement result, not a memory-only win

8. Anti-Patterns

Anti-pattern Why it fails Do this instead
“outputBackings means direct write” The property is explicitly a proposal Require returned-object identity and measured savings
“Raw malloc always copies on ANE” Public headers recommend page-aligned raw storage Measure raw and IOSurface arms on the target
“IOSurface means zero-copy” The API says it can avoid a copy, not that it always does Use conditional language and falsification
“One [4, ...] and one [2, ...] output” Those shapes belong to different model families Bind same-shaped K and V per family
“Batch supports the same backing contract” Batch is an explicit may-ignore case Prove synchronous single-example acceptance first
“Equal bytes prove the backing was used” Framework allocation plus copy can produce equal bytes Require object identity separately
“Metal buffer bridge is public” Installed MLFeatureValue.h exposes no such initializer Use public MLMultiArray or pixel-buffer paths
“Instruments has a definitive copy lane” No public guarantee covers every hidden copy Combine timing, allocation, identity, and parity evidence
“Deallocator runs on one safe background thread” No queue or thread is documented Make deallocation independent of thread assumptions
“Core ML serializes multiple output writes” Scheduling and alias semantics are not public Use distinct buffers and avoid overlap

9. Useful Hypotheses That Remain Unverified

The external report raised several ideas worth measuring, but not stating as platform facts:

  • an IOSurface-backed K/V pair may remove staging that remains for raw storage;
  • a persistent CVPixelBufferPool may reduce resident-loop allocation cost;
  • bounding temporary Objective-C objects with @autoreleasepool may reduce long-loop memory pressure;
  • post-processing with Accelerate over accepted bytes may reduce a later copy;
  • output backing may alter graph partitioning or compute-unit placement; and
  • multiple large outputs may have different acceptance or staging behavior from a single small output.

Each is a one-variable experiment. None justifies an arbitrary buffer registry, private API, a family-shape sweep, or a zero-copy product claim before target evidence exists.

Works Cited

  1. Apple: MLPredictionOptions.outputBackings — public proposal semantics, ignore cases, object-identity proof, backing restrictions, and page-alignment guidance; verified against the installed macOS 26.5 SDK MLPredictionOptions.h.
  2. Apple: MLMultiArray — external-pointer and pixel-buffer initializers, ownership, shapes, and scoped byte access; verified against installed MLMultiArray.h.
  3. Apple: MLFeatureDescription — feature constraints and isAllowedValue(_:); verified against installed MLFeatureDescription.h.
  4. Apple: MLFeatureValue — public value constructors; installed MLFeatureValue.h has no Metal-buffer constructor. — hypothesis source ingested with corrections; the absolute path and digest are frozen in Research Provenance because the sibling-relative link is workspace-local.