# External-Encoder, Decoder-Only Whisper Runtime on Apple Silicon

> **Collective Library edition.** This is the complete technical report. Private filesystem paths, internal run identifiers, campaign-control notes, and repository navigation were removed. Technical claims, code, measurements, evidence labels, citations, corrections, and falsification criteria are preserved.

August 12, 2026

## Executive Summary

Upstream whisper.cpp documents a hybrid configuration in which Core ML runs
the encoder while the decoder remains in the GGML model. It does **not**
document a decoder-only GGML container for this pinned loader. Structural
encoder deletion is a local experimental derivative and must be qualified as
such.

The opportunity is nevertheless concrete. The ordinary file initializer in
the pinned runtime opens the model with `std::ifstream`. The loader declares
the expected tensors, allocates backend buffers for them, then reads tensor
bytes into those allocations. It is not an `mmap`-backed model path. Enabling
Core ML changes encoder compute, but by itself does not delete the native
encoder tensors from the model allocation.

The smallest safe decoder-only design therefore has four parts:

1. derive an authenticated artifact that preserves the original header, mel
   filters, vocabulary, and every decoder record byte-for-byte while omitting
   every native encoder record;
2. make the runtime declare and expect only those retained decoder tensors;
3. make state creation and encoder execution fail closed unless the exact
   external Core ML sidecar loads and returns the required tensor; and
4. prove exact tokens, text, EOT behavior, memory deletion, and repeated-call
   lifecycle behavior against the complete-model control.

Do not mix this lossless structural experiment with further quantization,
cache shrinking, logits-API changes, or Core ML compression. Each changes a
different causal mechanism.

## 1. Corrected Evidence Boundary

| Raw-report claim | Corrected verdict |
| --- | --- |
| The pinned loader relies on a model `mmap` | **False.** The ordinary file path uses `std::ifstream`; the model loader allocates backend tensors and reads bytes into them. |
| Deleting encoder file bytes automatically saves the same resident bytes | **Not yet proven.** It prevents those tensors from being declared and loaded only when the runtime inventory changes too. Peak process and system memory still require measurement. |
| Decoder-only Whisper GGML is an upstream-documented format | **False.** Upstream documents a Core ML encoder paired with a GGML model. The stripped container and compatible loader are local experiments. |
| Structural deletion requires filtering a PyTorch checkpoint before conversion | **Too prescriptive.** A checked byte-for-byte derivative of an authenticated GGML source is simpler for the current experiment and avoids reconversion drift. |
| A caller must add `@autoreleasepool` or Core ML leaks about 5.8 MB per request | **Unsupported.** The checked bridge already wraps prediction in `@autoreleasepool`; the exact leak rate has no local or primary-source proof. Measure repeated-call growth before changing ownership. |
| `phys_footprint` alone predicts a macOS jetsam kill | **Unsupported as a singular rule.** It is a valuable task ledger, but desktop termination and product pressure also depend on working set, wired and clean pages, compression, swap, other processes, and system state. |
| `MLComputeUnits.all` proves ANE execution | **False.** It makes CPU, GPU, and Neural Engine available. Operation placement needs `MLComputePlan`, Instruments, or equivalent physical evidence. |
| Last-row logits and shortened K/V caches are mandatory for decoder-only correctness | **False.** They are separate mutable-state optimizations. First prove structural weight deletion with the baseline state and public API unchanged. |
| Q5 decoder output should exactly match an FP16 decoder | **Wrong control.** Structural deletion should exactly match the complete model with the same quantization. Q5-versus-FP16 quality is a separate numerical gate. |

## 2. Architecture and Ownership

The target runtime has one acoustic path and one text path:

```text
PCM
 |
 v
mel filters retained in GGML prefix
 |
 v
external Core ML encoder -----> [audio positions, audio state]
                                      |
                                      v
                           cross-attention K/V
                                      |
                                      v
                    retained GGML/Metal decoder
                                      |
                                      v
                             tokens and text
```

The Core ML package replaces native encoder **compute**. The decoder still
needs encoder hyperparameters because they define the external output and
cross-attention contract. Deleting native weights does not authorize changing
those dimensions.

### 2.1. Keep the ABI-bearing prefix

A decoder-only derivative should preserve, byte-for-byte:

- file magic and model hyperparameters;
- `n_audio_ctx`, `n_audio_state`, `n_audio_head`, and `n_audio_layer`;
- text dimensions and quantization metadata;
- mel-filter dimensions and values;
- vocabulary bytes and token IDs; and
- every retained decoder record header and payload.

Preserving `n_audio_layer` may look odd after encoder deletion. In this legacy
container it also participates in model-family identification and expected
architecture. The runtime patch, not an invented header value, should make
native encoder tensors unreachable.

### 2.2. Delete only native encoder records

Delete the records named by the pinned encoder inventory:

- positional embedding;
- two convolution weights and biases;
- final encoder layer normalization; and
- every tensor in every native encoder transformer block.

Retain all decoder self-attention and cross-attention tensors. Cross-attention
is part of the decoder even though its keys and values are derived from the
external encoder output.

### 2.3. One owner per resource

A memory-minimal serialized product starts with:

- one model context owning decoder weights;
- one state owning mutable caches and schedulers;
- one Core ML model handle associated with that state; and
- one request at a time through the context.

The public whisper.cpp contract does not make one context concurrently safe
for arbitrary state use. Add concurrency only after proving it is a product
requirement and measuring the incremental state and framework residency.

## 3. Deriving the Artifact

### 3.1. Prefer exact record filtering over reconversion

For a pinned, already-authenticated GGML source, the simplest inspected derivative
is a strict parser and copier:

1. parse the prefix through vocabulary exact EOF boundaries;
2. parse every tensor record with checked dimensions, type, name, byte count,
   and offset arithmetic;
3. classify each name against frozen encoder and decoder inventories;
4. copy the prefix and decoder record bytes without numerical conversion;
5. reject unknown, duplicate, missing, malformed, or trailing records; and
6. reparse the emitted artifact independently before sealing it.

This design isolates structural deletion. Filtering a PyTorch state dictionary
and rerunning conversion can also work, but introduces converter, source-model,
and quantizer drift that the deletion experiment does not need.

### 3.2. Seal both what remains and what disappeared

The derivation receipt should bind:

- complete source-model SHA-256 and byte size;
- prefix SHA-256 and byte size;
- ordered source tensor inventory;
- ordered retained and removed inventories;
- each retained record's source offset, byte length, type, shape, and hash;
- output SHA-256 and exact EOF;
- total retained and removed payload and physical record bytes;
- derivation script and focused test hashes; and
- the compatible runtime source, patch, build options, and binary closure.

An output hash alone cannot prove that the right tensors were retained. An
inventory alone cannot prove payload identity. Bind both.

### 3.3. Current 8e2d arithmetic is local experimental evidence

The current campaign's authenticated Q5 physics shell is 537,819,875 bytes.
An exact parse identifies 487 native encoder records and 52 decoder records.
Preserving its 611,057-byte prefix and all decoder records projects a stripped
artifact of 84,778,619 bytes, removing 453,041,256 physical file bytes.

Those numbers are **local binary-format evidence**, not upstream guarantees,
resident-memory measurements, quality evidence, or a shipping verdict. The
campaign note
owns their current receipts and status.

## 4. Making the Loader Decoder-Only

The pinned loader currently calculates tensor metadata, declares all native
encoder and decoder tensors, allocates backend buffers for them, and then
requires the number of loaded records to equal the declared tensor count.
Merely truncating the file therefore fails with missing tensors.

### 4.1. Change the expected inventory, not error tolerance

In an explicit external-encoder-only build or mode:

- do not create native encoder tensor metadata;
- do not add native encoder names to `model.tensors`;
- resize and create only the decoder-layer structures needed by text and
  cross-attention graphs;
- retain the stock unknown-tensor, shape, size, and loaded-count checks for the
  smaller inventory, then add exact type, duplicate, missing-name, and EOF
  checks; and
- log that native encoder capability is absent.

Do not make the generic loader accept arbitrary missing tensors. That converts
a precise derivative into a partially initialized model that can fail far from
the load boundary.

### 4.2. Make native encoder execution unreachable

Encoder pointers may remain null only if every path that can build or execute
the native encoder is structurally unavailable. The contract should require:

- Core ML compiled in;
- native fallback compiled out or rejected for this artifact class;
- Core ML sidecar load before a usable state is returned;
- an error return when prediction fails; and
- no attempt to build convolution or encoder graphs with missing tensors.

The model-only `_no_state` initializer may reasonably authenticate and load
decoder weights before a Core ML state exists. If so, document that boundary:
model-only load can succeed, while state creation and inference must fail when
the sidecar is absent. A negative control must exercise the exact public entry
point the product uses.

### 4.3. Validate the external encoder result

Do not treat a non-null `MLMultiArray` as sufficient. Check:

- prediction returned without error;
- output feature name is the expected one;
- element type is Float32 when the bridge copies Float32;
- rank and every dimension match the compiled contract;
- strides/layout match the copy path, or copy by checked strides; and
- element count exactly matches the destination allocation.

The external output shape and the preserved header's audio context must agree.
A shorter or longer output is a contract failure, not a reason to silently
truncate, pad stale memory, or invoke the absent native encoder.

### 4.4. The existing autorelease pool is not the missing feature

The checked Objective-C++ bridge already places the Core ML prediction call in
an `@autoreleasepool`. Under ARC, locally owned objects also leave scope with
the function. That does not prove the entire application is leak-free, but it
invalidates the raw report's claim that this bridge lacks a pool and leaks a
fixed 5.8 MB on every request.

If repeated inference grows memory, distinguish:

1. one-time Core ML specialization or allocator growth that plateaus;
2. reusable framework caching;
3. product-layer autoreleased objects outside the bridge; and
4. genuinely unbounded retained objects or buffers.

Use allocations and VM traces to identify the owner before adding another
pool or changing bridge lifetime.

## 5. Mutable State Is a Separate Optimization

Structural deletion targets immutable native encoder weights. Self-K/V,
cross-K/V, pad caches, schedulers, compute buffers, token vectors, and logits
remain. Their bounds depend on the product contract and should be changed only
after same-quantization parity passes.

### 5.1. Logits

The pinned state reserves capacity proportional to vocabulary size times text
context. A greedy product may ultimately need only current-step logits, but a
last-row-only change can alter public accessor assumptions, batching behavior,
and downstream pointer lengths. Measure committed memory first, then introduce
an explicit API contract and regression tests if the saving is material.

### 5.2. Text and audio caches

Bounding text length can reduce mutable state, but every caller must reject an
oversized request before graph construction. Bounding audio positions is not
valid merely because typical input is shorter: the external encoder output,
header metadata, cross cache, and decoder graph must all share the same frozen
maximum.

For the current full-span 8e2d experiment, preserve the 1,500-position external
encoder contract. Short-context state is a different artifact and quality
experiment.

### 5.3. Disable unused product features explicitly

Greedy, no-timestamp dictation can avoid beam candidates and optional DTW
state when the public product contract does not consume them. Record each
disabled feature in the receipt. Do not infer that a zero-valued request flag
prevented an earlier initialization allocation; inspect the construction
order.

## 6. Proving Memory Reduction on macOS

Artifact bytes, backend allocation logs, process footprint, and system pressure
answer different questions. No single column proves the whole result.

| Evidence | What it establishes | What it does not establish |
| --- | --- | --- |
| Artifact size and inventory | Native encoder records are absent on disk | Runtime allocation or Core ML residency |
| Backend weight-buffer log | Declared decoder tensors occupy a smaller backend buffer | Peak task or system memory |
| `TASK_VM_INFO.phys_footprint` and peak ledger | Memory charged to the worker task at sample times | Complete cross-process residency or universal termination risk |
| `/usr/bin/time -lp` maximum RSS and footprint | Coarse process peaks over a complete invocation | Exact phase attribution |
| `footprint` and `vmmap -summary` | Snapshot categories and VM-region attribution | A missed earlier peak |
| Core ML service census | Correlated work outside the app process | Causal attribution without a matched control |
| `vm_stat`, swap, and memory-pressure series | Host-level pressure during the experiment | Per-model ownership |
| `MLComputePlan` or Instruments placement | Eligible device use by operation on the tested build | Future devices, OS versions, or artifacts |

Sample at named phases:

1. before model load;
2. after decoder weights load;
3. after state and Core ML sidecar initialization;
4. after the first prediction and specialization;
5. across warmed repeated predictions; and
6. after state/context release and an idle observation window.

Compare complete-model and stripped-model processes with the same backend,
Core ML tree, input, request, thread count, warmup, host admission, and sample
cadence. Do not overlap the old and new contexts during replacement.

### 6.1. `ComputeUnits.all` is eligibility, not placement proof

Core ML Tools documents `ComputeUnits.ALL` as allowing all available units:
CPU, GPU, and Neural Engine. The same documentation exposes
`MLComputePlan.get_compute_device_usage_for_mlprogram_operation` to inspect
planned operation use. Therefore a successful load or prediction under `ALL`
does not, by itself, prove ANE residency.

For a physical claim, bind the compiled-tree hash and capture an operation-level
compute plan or Instruments trace on the exact target host. Report boundary
CPU/GPU operations rather than collapsing a mixed graph into “runs on ANE.”

## 7. Correctness and Negative Controls

The structural-deletion control must use the same quantized decoder and same
external Core ML encoder on both sides. Require:

- identical token IDs in order;
- identical EOT presence and stop reason;
- identical decoded-text bytes;
- identical token ceilings and request options;
- deterministic output across repetitions; and
- no latency regression beyond a preregistered tolerance.

Include full-span inputs near both 15 and 30 seconds. A load-only pass cannot
prove that null native-encoder pointers stay unreachable during real
prediction.

Named negative controls should reject:

- missing or renamed Core ML sidecar;
- Core ML prediction error or wrong output shape/type;
- one missing decoder record;
- one duplicate decoder record;
- an unknown encoder record reintroduced into the stripped artifact;
- wrong tensor type, dimensions, or payload length;
- truncated input and trailing bytes;
- native-fallback configuration; and
- more tokens or audio positions than the bounded state allows.

Q5-versus-FP16 WER and pathology qualification remains separate. Exact
stripped-Q5 versus complete-Q5 parity proves deletion is lossless; it does not
prove Q5 is good enough for production.

## 8. Failure Modes and Debugging Playbook

### 8.1. Loader reports an unknown or missing tensor

The derivative inventory and runtime declaration disagree. Print the first
unexpected name, dimensions, type, record offset, and expected inventory hash.
Do not suppress the error or permit partial loading.

### 8.2. Load succeeds, then native graph construction crashes

A fallback or graph-builder path still reaches a deleted encoder pointer.
Reproduce with the Core ML sidecar missing, then make the external-encoder-only
mode fail before any inference graph can be built.

### 8.3. Output repeats or hallucinates after the split

First compare the external encoder output shape and bytes, cross-cache
construction, request reset, and exact decoder artifact against the complete
control. Do not attribute the symptom to encoder deletion until those contracts
match. Stale cross-K/V and audio-context mismatch can produce plausible text
without a clean crash.

### 8.4. The file shrinks but peak memory does not

Check the model-load buffer log. If it still reports the complete weight size,
the runtime still declares native tensors. If that buffer shrank, split the
remaining peak among Core ML initialization, caches, compute arenas, temporary
read buffers, dynamic-library load, and overlapping contexts.

### 8.5. Memory grows across repeated predictions

Run enough iterations to distinguish lazy plateau from linear growth. Mark
pool boundaries, model/state lifetime, and request completion in the trace.
Because the bridge already has a prediction-local autorelease pool, inspect
actual retained owners before diagnosing an autorelease leak.

### 8.6. Core ML loads, but ANE evidence is absent

This is expected under a permissive compute-unit request. Capture the exact
compiled model's compute plan or Instruments trace. Treat load logs as Core ML
utilization evidence only, not device-placement evidence.

## 9. Anti-Patterns

| Anti-pattern | Why it fails | Do this instead |
| --- | --- | --- |
| Delete bytes and relax “not all tensors loaded” | Accepts partially initialized arbitrary models | Declare an exact smaller inventory and retain strict EOF/count checks |
| Change header audio dimensions to describe missing weights | Breaks external output and cross-attention sizing | Preserve ABI-bearing metadata; remove native capability in code |
| Keep native fallback “for safety” | Deleted tensors make fallback impossible | Fail state creation or prediction closed |
| Reconvert and requantize during the deletion control | Mixes structural and numerical changes | Copy authenticated decoder records byte-for-byte |
| Claim the disk-byte delta as RAM saved | Ignores state, services, allocator behavior, and peaks | Run matched process and system measurements |
| Add another autorelease pool from a fixed leak anecdote | Changes lifetime without identifying an owner | Prove plateau versus leak with traces |
| Claim ANE from `ComputeUnits.all` or a load log | Availability and Core ML use do not identify placement | Bind `MLComputePlan` or Instruments evidence |
| Shrink caches and logits in the first patch | Makes a lossless parity failure ambiguous | Prove weight deletion first, optimize state second |
| Compare stripped Q5 directly with FP16 for exactness | Quantization itself can change tokens | Compare stripped and complete models at the same quantization |

## 10. Decision Framework

| Approach | Value | Main risk | Use when |
| --- | --- | --- | --- |
| Stock hybrid model | Upstream-documented control | Retains unused native encoder allocation | Establishing correctness and memory baseline |
| Full-file validate-and-skip loader | Isolates allocation deletion without changing file bytes | More runtime parsing/skip machinery; forbidden bytes remain shippable | Temporary losslessness falsifier |
| Authenticated stripped artifact plus explicit loader | Smallest disk and declared weight inventory | Local format/runtime fork needs strong receipts | Preferred experiment after mechanics are understood |
| Reconverted decoder-only checkpoint | Can integrate into a future producer | Adds conversion and quantization drift | Only when source-level production generation is required |
| Mutable-state reduction | Additional RAM after weights | API, capacity, and correctness regressions | After structural parity and measured state attribution |
| Core ML compression | May reduce package and residency | Numerical drift and placement changes | Last, under a separate physical qualification |

The clean promotion order is:

1. complete-model same-quantization control;
2. stripped artifact and fail-closed loader;
3. exact inference parity and malformed-input controls;
4. matched local physical-memory proof;
5. exact target-Mac latency, placement, memory, and swap proof;
6. mutable-state experiments one mechanism at a time; and
7. independent decoder quantization and product-quality qualification.

## Works Cited

1. [whisper.cpp pinned source tree][whisper-pin] — source pin used for the
   legacy loader, tensor, state, and Core ML boundary.
2. [whisper.cpp Core ML support][whisper-coreml] — upstream hybrid encoder
   setup and runtime load-log contract.
3. [Core ML Tools compute-unit options][coreml-compute-units] — `ALL`, CPU,
   GPU, and Neural Engine eligibility choices.
4. [Core ML Tools model utilities][coreml-model-utilities] — `MLComputePlan`
   operation-device and estimated-cost inspection.
5. Low-memory Whisper Turbo inference —
   locally corrected Darwin accounting, pinned allocation topology, and memory
   receipt design.
6. [Physical Core ML plus whisper.cpp verification][physical-verification] —
   local fail-closed and placement-evidence contract.

[whisper-pin]: https://github.com/ggerganov/whisper.cpp/tree/97c56f1dc1d1100a9d859c865a20c82d22f823ed
[whisper-coreml]: https://github.com/ggerganov/whisper.cpp#core-ml-support
[coreml-compute-units]: https://apple.github.io/coremltools/docs-guides/source/load-and-convert-model.html#set-the-compute-units
[coreml-model-utilities]: https://apple.github.io/coremltools/docs-guides/source/mlmodel-utilities.html
[physical-verification]: /reports/whispercpp-coreml-physical-verification
