# Core ML Preallocated Multi-Output Buffers on Apple Silicon

> **Collective Library edition.** This is the complete technical report. Private filesystem paths, internal run identifiers, campaign-control notes, and repository navigation were removed. Technical claims, code, measurements, evidence labels, citations, corrections, and falsification criteria are preserved.

August 9, 2026

## Executive Summary

`MLPredictionOptions.outputBackings` is a public way to **propose** destination
objects for Core ML outputs. It is not a direct-write command. Core ML may
ignore a proposed backing when the model does not support it or when prediction
uses batch mode, and an unknown output name is silently ignored. A successful
prediction is therefore insufficient evidence. The caller must prove that the
returned output is the exact proposed object.

For a multi-array output, the backing must be an `MLMultiArray` accepted by the
output's `MLFeatureDescription.isAllowedValue(_:)`. Apple explicitly recommends
a page-aligned address for a raw external-memory `MLMultiArray`. The public
initializer wraps that storage without copying when the wrapper is created.
Neither fact proves that the selected compute device writes directly into those
bytes; Core ML may still stage internally.

An IOSurface-backed `CVPixelBuffer`, wrapped as an `MLMultiArray`, can reduce
inference latency by avoiding a buffer copy. That is the SDK's deliberately
conditional wording. It is not a blanket zero-copy guarantee. There is no
public `MLFeatureValue` initializer over `MTLBuffer` in the installed SDK.

For one Whisper family invocation, the two outputs are K and V with the same
shape. Turbo uses two `[4, 1536, 1280]` FP16 outputs; Distil uses two
`[2, 1536, 1280]` FP16 outputs. These are alternative packages. A single
invocation never combines one Turbo-shaped output with one Distil-shaped
output.

Treat direct output ownership as a falsifiable optimization:

1. validate names, feature types, shapes, data types, and strides;
2. run one synchronous prediction with distinct K and V backings;
3. require returned-object identity for both outputs;
4. prove numerical parity, complete writes, guard preservation, and zero tails;
5. measure allocation, packing, copying, and end-to-end warmed latency; and
6. keep placement evidence separate from backing acceptance.

## 1. The Contract: Proposal, Validation, and Proof

### 1.1. Public API and version matrix

| API or behavior | Availability in installed SDK | What is documented | What is not documented |
| --- | --- | --- | --- |
| `MLPredictionOptions.outputBackings` | macOS 11, iOS 16, watchOS 9, tvOS 16 | Dictionary from feature name to `MLMultiArray` or `CVPixelBuffer`, according to feature type | Acceptance by every `mlprogram`, direct device writes, or a latency saving |
| `MLMultiArray.init(dataPointer:shape:dataType:strides:deallocator:)` | Public `MLMultiArray` API | Wraps existing storage without a copy at construction | No internal staging during prediction or a documented deallocator queue |
| `MLMultiArray.init(pixelBuffer:shape:)` | macOS 12, iOS 16, watchOS 9, tvOS 16 | Wraps and owns an IOSurface-backed pixel buffer; can reduce latency by avoiding a copy | Zero copies for every model, compute unit, OS, shape, or prediction mode |
| `MLFeatureDescription.isAllowedValue(_:)` | Public Core ML API | Validates a feature value against its declared output description | That passing validation forces Core ML to use the backing |
| Batch prediction | Public Core ML API | `outputBackings` may be ignored in batch mode | Stable object reuse or caller-controlled output allocation for a batch |
| `MLFeatureValue` over `MTLBuffer` | No public initializer in installed SDK | Nothing | Direct public Metal-buffer binding to a model output |

The `outputBackings` header defines five fail-closed rules:

- a missing feature entry permits framework allocation;
- an unsupported model permits framework allocation;
- batch prediction permits framework allocation;
- an unknown feature name is ignored; and
- a known backing must satisfy the output feature description's
  `isAllowedValue(_:)` test or prediction reports an error.

After prediction, compare the returned multi-array to the proposed backing by
**object identity**. Equal shapes, equal pointers observed at different times,
or equal bytes are weaker claims. They do not prove that Core ML accepted the
proposed object.

### 1.2. What “without copy” actually establishes

The raw-pointer initializer's “without copy” statement is about constructing
the `MLMultiArray` wrapper around existing storage. It establishes that the
wrapper does not clone the initial bytes. It does not specify how CPU, GPU, or
Neural Engine execution reaches that storage.

Likewise, the pixel-buffer initializer says an IOSurface-backed array *can*
reduce inference latency by avoiding a buffer copy. It does not promise that
every output is written in place, that no temporary exists, or that an accepted
backing preserves Neural Engine placement.

The categorical claim that the Neural Engine always copies to plain `malloc`
storage is therefore unsupported. It also conflicts with Apple's explicit
recommendation to use a page-aligned raw `MLMultiArray` address for best
`outputBackings` performance. Test both mechanisms on the target model.

## 2. Family-Specific Fixed Shapes

The shape contract must be derived from one model package, not assembled across
families.

| Family package | K shape | V shape | C-order element strides | Bytes per output | K plus V |
| --- | --- | --- | --- | ---: | ---: |
| Turbo | `[4, 1536, 1280]` | `[4, 1536, 1280]` | `[1966080, 1280, 1]` | 15,728,640 | 30 MiB |
| Distil | `[2, 1536, 1280]` | `[2, 1536, 1280]` | `[1966080, 1280, 1]` | 7,864,320 | 15 MiB |

The first stride is independent of layer count because one layer spans
`1536 * 1280 = 1,966,080` elements. FP16 uses two bytes per element. The
runtime contract must compare the actual output description and returned array
against these values; arithmetic in a guide is not an ABI.

The deployed cross-attention cache uses only positions `0:1500`. Positions
`1500:1536` are physical padding and must remain bitwise FP16 zero in the
campaign candidate. A zeroed destination before prediction is not sufficient:
the post-prediction verifier must inspect every tail byte for K and V.

### 2.1. Pixel-buffer layout

For `MLMultiArray.init(pixelBuffer:shape:)`, the installed header documents an
exact mapping: the last shape dimension equals pixel-buffer width, and the
product of all preceding dimensions equals height. Therefore a Turbo backing
would use width 1280 and height 6144; a Distil backing would use width 1280 and
height 3072.

The pixel format must be `kCVPixelFormatType_OneComponent16Half` for FP16. A
consumer must still respect `CVPixelBufferGetBytesPerRow`; Core Video may pad
rows. Do not turn the documented shape mapping into an assumption of dense row
bytes.

## 3. Buffer Mechanisms

| Mechanism | Documented fact | Working interpretation |
| --- | --- | --- |
| Page-aligned raw storage plus `MLMultiArray` | Apple recommends page alignment; initializer wraps existing storage without copying | Simplest C/C++ ownership path and the first mechanism to test |
| IOSurface-backed `CVPixelBuffer` wrapped as `MLMultiArray` | Can reduce latency by avoiding a copy; wrapper owns the pixel buffer | Strong alternative when raw storage is ignored or slower, but still requires identity and timing proof |
| `CVPixelBufferPool` | Apple calls an IOSurface-backed pool often efficient, especially for playback or export | May reduce repeated allocation overhead; relevance to large tensor outputs is unverified |
| Framework-allocated `MLMultiArray` | Always available as the control | Safest ownership baseline, with an explicit checked copy into runtime cache |
| `MTLBuffer` as `MLFeatureValue` | No public initializer exists in installed `MLFeatureValue.h` | Do not ship a private or invented bridge |
| Accelerate over returned bytes | Accelerate can process caller-accessible memory after prediction | Useful post-processing, but not evidence that prediction was zero-copy |

IOSurface can reduce copies across compatible framework boundaries. It does not
make “zero-copy” an API property. Promote it only after the exact target model,
OS, compute units, and prediction path prove identity, parity, and a measurable
end-to-end saving.

## 4. Ownership, Lifetime, and Concurrency

| Object | Caller obligation | Public guarantee boundary |
| --- | --- | --- |
| Raw allocation | Keep bytes valid until the wrapper is deallocated | Deallocator is called when the array is deallocated; its thread is unspecified |
| `MLMultiArray` backing | Keep the object alive through prediction and result validation | Core ML may return a different array |
| Pixel buffer | Do not lock its base address during prediction | Pixel-buffer-backed array owns and releases the buffer |
| K and V backings | Use distinct storage and validate both returned identities | No public output-write ordering or aliasing guarantee |
| Reused backing | Serialize access until the prior synchronous prediction and reads finish | No public guarantee for overlapping predictions into one object |

The installed headers do not promise a queue or thread for the external-memory
deallocator. Write the closure so it has no main-thread requirement, captures
only stable ownership, and safely releases exactly once. Do not state that Core
ML uses one background deallocation thread.

The public API also does not specify whether multiple outputs are produced
serially or concurrently. Never alias K and V, never alias an input and output,
and do not overlap predictions that reuse the same storage. These rules are
conservative caller policy, not claims about Core ML's hidden scheduler.

An `@autoreleasepool` around a long-running Objective-C loop can bound
temporary Objective-C objects. Community reports associate unbounded loops
with IOSurface pressure, but the installed headers do not make one
autoreleasepool per prediction mandatory or guarantee that it cures every pool
exhaustion failure. Treat it as an operational experiment with resident-memory
telemetry.

## 5. Minimal Implementation Patterns

### 5.1. Objective-C++ page-aligned raw backings

This sketch uses the SDK-recommended raw-memory path. It intentionally does not
call it a direct Neural Engine write.

```objc
#import <CoreML/CoreML.h>
#import <mach/vm_page_size.h>

static MLMultiArray *MakeFP16Backing(NSArray<NSNumber *> *shape,
                                      size_t elementCount,
                                      NSError **error) {
    void *bytes = NULL;
    const size_t byteCount = elementCount * sizeof(uint16_t);
    if (posix_memalign(&bytes, vm_page_size, byteCount) != 0) {
        return nil;
    }

    MLMultiArray *array = [[MLMultiArray alloc]
        initWithDataPointer:bytes
                      shape:shape
                   dataType:MLMultiArrayDataTypeFloat16
                    strides:@[@1966080, @1280, @1]
                deallocator:^(void *ownedBytes) { free(ownedBytes); }
                      error:error];
    if (array == nil) {
        free(bytes);
    }
    return array;
}

NSArray<NSNumber *> *shape = @[@4, @1536, @1280]; // Turbo only.
const size_t elementCount = 4ULL * 1536ULL * 1280ULL;
MLMultiArray *crossK = MakeFP16Backing(shape, elementCount, &error);
MLMultiArray *crossV = MakeFP16Backing(shape, elementCount, &error);

MLFeatureDescription *kDescription =
    model.modelDescription.outputDescriptionsByName[@"cross_k"];
MLFeatureDescription *vDescription =
    model.modelDescription.outputDescriptionsByName[@"cross_v"];

BOOL kAllowed = [kDescription
    isAllowedValue:[MLFeatureValue featureValueWithMultiArray:crossK]];
BOOL vAllowed = [vDescription
    isAllowedValue:[MLFeatureValue featureValueWithMultiArray:crossV]];
if (!kAllowed || !vAllowed) {
    // Reject the candidate before prediction; do not infer acceptance.
}

MLPredictionOptions *options = [[MLPredictionOptions alloc] init];
options.outputBackings = @{@"cross_k": crossK, @"cross_v": crossV};
id<MLFeatureProvider> result =
    [model predictionFromFeatures:input options:options error:&error];

MLMultiArray *returnedK =
    [result featureValueForName:@"cross_k"].multiArrayValue;
MLMultiArray *returnedV =
    [result featureValueForName:@"cross_v"].multiArrayValue;
if (returnedK != crossK || returnedV != crossV) {
    // Core ML did not accept both proposals. Use the checked-copy control.
}
```

Production code should centralize the family shape and byte count, check every
allocation and feature lookup, retain the backings for the full use interval,
and make failure select an explicit control path. Never continue as if an
ignored backing were accepted.

### 5.2. Swift identity check

The same proof is required from Swift:

```swift
let options = MLPredictionOptions()
options.outputBackings = ["cross_k": crossK, "cross_v": crossV]

let result = try model.prediction(from: input, options: options)
guard result.featureValue(for: "cross_k")?.multiArrayValue === crossK,
      result.featureValue(for: "cross_v")?.multiArrayValue === crossV else {
    throw BackingError.outputBackingIgnored
}
```

Do not use batch prediction for the first acceptance proof. The public header
explicitly names batch as a mode where the proposed backing may not be used.

### 5.3. IOSurface-backed alternative

When testing the pixel-buffer path:

1. create an IOSurface-backed `CVPixelBuffer` with the exact width, height, and
   `kCVPixelFormatType_OneComponent16Half` format;
2. wrap it with `MLMultiArray.init(pixelBuffer:shape:)`;
3. do not lock the base address or call array data/subscript accessors before
   prediction;
4. propose the wrapper through `outputBackings`;
5. require returned `MLMultiArray` identity; and
6. lock and inspect bytes only after synchronous prediction completes.

The wrapper owns the pixel buffer. Audit retains and releases accordingly; do
not release storage as though ownership remained solely with the creator.

## 6. Falsification Recipe

Separate five questions that are easy to blur together:

| Question | Required evidence |
| --- | --- |
| Is the API available? | Compile against the target SDK and record deployment target |
| Is this value admissible? | Exact output name plus `isAllowedValue(_:)` success |
| Was this object used? | Returned K and V object identity |
| Is the result correct? | Shape, type, stride, range, finite-value, parity, guard, digest, and tail checks |
| Did the product get faster? | Paired warmed stage and WAV-to-final-text timings under a quiet-host gate |

### 6.1. Acceptance and corruption checks

For every scored prediction:

- require the exact model output-name set before constructing options;
- require the expected feature type, shape, data type, and element strides;
- require `isAllowedValue(_:)` for each proposed backing;
- require returned-object identity for both K and V;
- require the returned pointer range to stay within the owned allocation;
- compare all logical elements against the framework-allocation control under a
  frozen numerical tolerance;
- hash all logical output bytes, not a prefix sample;
- place canaries around caller-owned storage and require them unchanged;
- require every padded tail byte to remain zero; and
- reject NaN, infinity, all-zero logical tensors, stale prior-call digests, and
  K/V aliasing.

Object identity proves that Core ML returned the proposed object. It does not
prove there was no hidden staging before the result arrived. Latency and
allocation evidence address that separate hypothesis.

### 6.2. Timing and allocation A/B

Compare exactly three arms with the same compiled model and inputs:

1. framework output plus explicit checked copy;
2. page-aligned raw `MLMultiArray` output backing; and
3. IOSurface-backed `MLMultiArray` output backing.

Keep model load, warmup count, compute units, thread policy, input order, and
host state fixed. Report prediction, packing/copy, complete encoder stage, and
WAV-to-final-text time separately. Include allocation and wrapper creation in
the interval unless the product reuses those objects persistently.

Use allocation tracing, Time Profiler, signposts around the owned stages, and a
compute-plan inspection as supporting evidence. The public SDK does not
guarantee an Instruments lane that exposes every Core ML copy, nor does absence
of a visible copy event prove zero-copy execution. Static compute-unit choice
also does not prove Neural Engine residency or specify fallback behavior.

### 6.3. Decision rules

- **Reject API feasibility** if either output name is absent, either proposed
  value is disallowed, or either returned identity differs.
- **Reject correctness** on any parity, range, stride, canary, tail, stale-data,
  or termination failure.
- **Reject copy-removal language** unless the accepted backing produces a
  repeatable paired stage saving beyond timer noise and allocation effects.
- **Reject product promotion** unless the saving survives the campaign's full
  end-to-end latency and quality gates.
- **Retain the explicit-copy control** whenever the runtime, OS, compiled model,
  family, or output description changes.

## 7. Failure Modes

| Symptom | Likely cause | Fail-closed response |
| --- | --- | --- |
| Prediction succeeds but identity differs | Unsupported backing, batch mode, or omitted entry | Score as ignored; use checked copy |
| Misspelled output appears harmless | Unknown names are ignored | Compare exact output-name sets before prediction |
| Prediction reports an admissibility error | Backing fails feature description | Reject shape, type, or storage choice; do not coerce silently |
| Pixel-buffer path stalls or fails | Base address or array data access locked it before prediction | Recreate unlocked backing and enforce access order |
| Correct shape but corrupted cache | Wrong element strides or ignored row bytes | Validate strides and use `CVPixelBufferGetBytesPerRow` |
| Tail contains nonzero values | Model/package wrote padding or verifier inspected wrong range | Reject artifact; never mask after the fact |
| Repeated calls return stale bytes | Overlap, lifetime error, or partial write | Serialize, seed distinct canaries/digests, and reject repeats |
| Candidate allocates less but is not faster | Allocation was not the bottleneck or staging remains | Keep control and stop optimization |
| Placement changes with backing | Runtime selected a different execution plan | Treat as a separate placement result, not a memory-only win |

## 8. Anti-Patterns

| Anti-pattern | Why it fails | Do this instead |
| --- | --- | --- |
| “`outputBackings` means direct write” | The property is explicitly a proposal | Require returned-object identity and measured savings |
| “Raw `malloc` always copies on ANE” | Public headers recommend page-aligned raw storage | Measure raw and IOSurface arms on the target |
| “IOSurface means zero-copy” | The API says it can avoid a copy, not that it always does | Use conditional language and falsification |
| “One `[4, ...]` and one `[2, ...]` output” | Those shapes belong to different model families | Bind same-shaped K and V per family |
| “Batch supports the same backing contract” | Batch is an explicit may-ignore case | Prove synchronous single-example acceptance first |
| “Equal bytes prove the backing was used” | Framework allocation plus copy can produce equal bytes | Require object identity separately |
| “Metal buffer bridge is public” | Installed `MLFeatureValue.h` exposes no such initializer | Use public `MLMultiArray` or pixel-buffer paths |
| “Instruments has a definitive copy lane” | No public guarantee covers every hidden copy | Combine timing, allocation, identity, and parity evidence |
| “Deallocator runs on one safe background thread” | No queue or thread is documented | Make deallocation independent of thread assumptions |
| “Core ML serializes multiple output writes” | Scheduling and alias semantics are not public | Use distinct buffers and avoid overlap |

## 9. Useful Hypotheses That Remain Unverified

The external report raised several ideas worth measuring, but not stating as
platform facts:

- an IOSurface-backed K/V pair may remove staging that remains for raw storage;
- a persistent `CVPixelBufferPool` may reduce resident-loop allocation cost;
- bounding temporary Objective-C objects with `@autoreleasepool` may reduce
  long-loop memory pressure;
- post-processing with Accelerate over accepted bytes may reduce a later copy;
- output backing may alter graph partitioning or compute-unit placement; and
- multiple large outputs may have different acceptance or staging behavior
  from a single small output.

Each is a one-variable experiment. None justifies an arbitrary buffer registry,
private API, a family-shape sweep, or a zero-copy product claim before target
evidence exists.

## Works Cited

1. [Apple: `MLPredictionOptions.outputBackings`](https://developer.apple.com/documentation/coreml/mlpredictionoptions/outputbackings)
   — public proposal semantics, ignore cases, object-identity proof, backing
   restrictions, and page-alignment guidance; verified against the installed
   macOS 26.5 SDK `MLPredictionOptions.h`.
2. [Apple: `MLMultiArray`](https://developer.apple.com/documentation/coreml/mlmultiarray)
   — external-pointer and pixel-buffer initializers, ownership, shapes, and
   scoped byte access; verified against installed `MLMultiArray.h`.
3. [Apple: `MLFeatureDescription`](https://developer.apple.com/documentation/coreml/mlfeaturedescription)
   — feature constraints and `isAllowedValue(_:)`; verified against installed
   `MLFeatureDescription.h`.
4. [Apple: `MLFeatureValue`](https://developer.apple.com/documentation/coreml/mlfeaturevalue)
   — public value constructors; installed `MLFeatureValue.h` has no Metal-buffer
   constructor.
   — hypothesis source ingested with corrections; the absolute path and digest
   are frozen in Research Provenance because the sibling-relative link is
   workspace-local.
