# Stateful Core ML Whisper Decoders on Apple Silicon

> **Collective Library edition.** This is the complete technical report. Private filesystem paths, internal run identifiers, campaign-control notes, and repository navigation were removed. Technical claims, code, measurements, evidence labels, citations, corrections, and falsification criteria are preserved.

August 9, 2026

This guide separates the documented Core ML state API from attractive but
unproved claims about Neural Engine residency and base-M1 speed. It is a field
guide for falsifying an internal-cache Whisper decoder, not evidence that one
is faster than the existing CPU decoder.

See also the August 2026 external-research refresh, which preserves public implementation recipes while annotating them with this guide's local placement falsifier.

The raw report was corrected during ingest against Core ML Tools 9, the local
macOS 15 SDK headers, the public WhisperKit paper, and pinned public
WhisperKitTools source. Claims about cache traffic, operation placement, and
M1 latency remain hypotheses until measured.

## 1. Evidence Boundary

Four statements are documented:

1. Stateful `mlprogram` models are supported from iOS 18 and macOS 15.
2. `ct.StateType` wraps a tensor state during conversion.
3. A runtime model can allocate an `MLState` and pass it to prediction.
4. Core ML Tools can represent state mutation with `read_state` and
   `coreml_update_state`; graphs may contain `slice_update`.

None of those statements proves that a Whisper-shaped cache stays on the ANE,
that a one-slot update avoids full-state traffic, or that the resulting token
loop beats whisper.cpp on a base M1.

The WhisperKit paper reports a real component result: the authors changed the
large-v3-turbo decoder from cache-as-input tensors to stateful Core ML and
reported a forward-pass reduction from 8.4 to 4.6 ms on an M3 ANE, or 45%.
The paper does not provide the token count, timing distribution, base-M1
result, compiled package, or enough source identity to reproduce that number.

Local source correction: the pinned WhisperKitTools checkout at commit
`84f77a83c8f530022ae55fbb1a64b3351ef63c7a` contains no `StateType` or
`MLState`. Its decoder accepts full self-key/value caches as ordinary inputs,
returns one-token cache updates, and accepts encoder embeddings on every token.
Its pipeline defaults the decoder to `cpuAndGPU`. Therefore the paper is prior
art, but this public checkout is not an executable MLState reference for the
paper's 4.6 ms result.

## 2. Exact Runtime Contract

The macOS 15 SDK describes `MLState` as a handle to model-owned buffers. The
important lifecycle rules are:

- `newState()` allocates the model's declared state buffers and initializes
  them to zero.
- `prediction(from:using:)` and its asynchronous equivalent consume a state
  handle.
- Predictions sharing one state must be serialized. Reading or writing a state
  while a prediction uses it is undefined behavior.
- `withMultiArray` exposes mutable initialization/access, but the buffer address
  may differ across closures. Never retain a pointer beyond the closure.
- There is no semantic reset operation. Allocate a fresh zeroed state for each
  independent utterance or decode branch.

One state per independent request is the simple safe rule. Parallel token
calls on one state are invalid; parallel requests require independent states.

```swift
let configuration = MLModelConfiguration()
configuration.computeUnits = .cpuAndNeuralEngine

let model = try MLModel(contentsOf: modelURL, configuration: configuration)
let state = model.makeState()

// Serialize every prediction that shares `state`.
let output = try await model.prediction(from: input, using: state)
```

Requested `.cpuAndNeuralEngine` is a scheduler constraint, not placement
evidence. Preserve a `.cpuOnly` negative control and inspect the compiled
compute plan.

## 3. Conversion Contract

Core ML Tools 9 requires `StateType.wrapped_type` to be a `TensorType`. The
wrapped tensor cannot carry its own name or default value. The source state
name must match the traced model buffer name.

```python
import coremltools as ct
import numpy as np

states = [
    ct.StateType(
        wrapped_type=ct.TensorType(
            shape=(4, 20, 448, 64),
            dtype=np.float16,
        ),
        name="self_k",
    ),
]

model = ct.convert(
    traced_model,
    inputs=inputs,
    states=states,
    convert_to="mlprogram",
    minimum_deployment_target=ct.target.macOS15,
)
```

For the pinned stack, treat state tensors as fixed-shape FP16 buffers. Do not
generalize the raw report's statements that every dynamic input or every batch
larger than one necessarily falls back; those claims were not established by
the cited API. Falsify the exact shape needed by the decoder.

### State update is a question, not an answer

The expected MIL shape is conceptually:

```text
read_state -> slice_update(active position) -> coreml_update_state
```

That graph proves semantics, not traffic. A compiler may lower a one-slot
update efficiently, materialize a larger tensor, or partition adjacent ops.
Measure latency while sweeping cache length and update position. If latency
scales with the full state size, the mechanism has failed even if outputs are
correct.

## 4. Whisper Cache Arithmetic

For large-v3-turbo, the decoder has four layers, 20 heads, head dimension 64,
and width 1280. With a 448-token self-cache and FP16 K/V tensors:

```text
4 layers * 2 K/V * 20 heads * 448 positions * 64 values * 2 bytes
= 9,175,040 bytes
= 8.75 MiB
```

Cross-attention K/V for the full 1500 encoder positions is:

```text
4 layers * 2 K/V * 20 heads * 1500 positions * 64 values * 2 bytes
= 30,720,000 bytes
= 29.30 MiB
```

The combined turbo cache is 38.05 MiB. Distil-large-v3.5 has two decoder
layers, so its corresponding cache is half that size. This arithmetic does not
show how many bytes the runtime moves per token.

The current turbo FP16 decoder weights are 343,818,240 bytes. At the base M1's
approximate 68 GB/s memory bandwidth, one ideal stream of those weights already
costs about 5.06 ms before cache reads, compute, dispatch, state update, or
output handling. A 4.6 ms M3 result therefore cannot be projected directly to
base M1.

## 5. Graph Partitioning

There are three useful experimental boundaries:

| Boundary | Question answered | What it does not prove |
| --- | --- | --- |
| State-only update | Does update cost scale with cache size? | Decoder placement or quality |
| One decoder layer | Do attention and state mutation lower together? | Full head and package behavior |
| Full decoder | Does the product token loop win and remain exact? | Encoder or end-to-end target |

A multifunction package can combine `prefill` and `decode` functions and
deduplicate byte-identical constants. Core ML Tools documents this capability;
it does not guarantee shared runtime residency or that two functions can mutate
and consume one state exactly as desired. Prove the actual state handoff.

Separate Turbo and Distil artifacts are simplest. They have different decoder
depths and learned weights. A single dynamic supergraph adds a branch and does
not create useful learned-weight sharing.

### Prefill and cross-attention

The attractive architecture is:

```text
encoder output
    -> prefill/setup function
        -> cross K/V state
        -> prompt self K/V state
    -> serialized one-token decode calls
        -> updated self K/V state
        -> next-token decision
```

This is not yet a documented multifunction-state recipe. Test whether the same
state object can be initialized through one function and consumed through
another before designing the full decoder around it.

Whisper.cpp already owns both `kv_self` and read-only `kv_cross`. Statefulness
alone deletes no decoder math. The prize comes only if Core ML reduces actual
cache/weight traffic or executes the decoder body faster.

## 6. The Vocabulary Head Is a Separate Risk

The decoder body and the 51,866-way output projection have different placement
and bandwidth risks. Measure these variants explicitly:

1. decoder body with a small diagnostic output;
2. monolithic full-vocabulary head;
3. four fixed output-head shards below suspected compiler dimension cliffs;
4. in-graph top-k or argmax only.

Returning only argmax is insufficient for Whisper unless the graph implements
the exact suppression, no-speech, timestamp, EOT, and fallback-temperature
policy owned by the product contract. For the current deterministic final-text
campaign, compare the complete greedy token sequence, not merely top-1 logits
on isolated steps.

Embedding gather placement and tied embedding/head storage must also be
inspected. A compiled package may duplicate constants even when the source
model ties them.

## 7. Minimal Physical Falsifier Ladder

Run these in order on the exact base M1/macOS 15 target:

### 7.1. State-only bandwidth probe

- Allocate exact self and cross-cache shapes.
- Compare ordinary tensor inputs with `MLState`.
- Sweep active cache length and update position.
- Include fresh-state, repeated-call, and two-independent-state controls.
- Record p50, p95, first-run specialization, output hashes, and compute plan.

Kill if one-slot latency scales like full-cache materialization, state reset is
nondeterministic, or independent states contaminate one another.

### 7.2. Exact one-layer decoder

- Width 1280, 20 heads, query length one.
- Use the exact normalization, residual, self-attention, and cross-attention
  operations from the candidate model.
- Compare explicit cache inputs against MLState with identical weights.
- Omit the vocabulary head so placement failures are attributable.

Kill if the stateful form is not faster than the explicit form or introduces a
material CPU/GPU island under `CPU_AND_NE`.

### 7.3. Full body plus head ladder

- Scale to two layers for Distil and four for Turbo.
- Add the head variants separately.
- Run forced-prefix-to-EOT parity against the pinned CPU decoder.
- Prove cache reset across utterances and deterministic termination.

### 7.4. Resident end-to-end A/B

The sealed CPU decoder baseline spent 203.3455 ms on 28 decode evaluations,
or 7.2623 ms/evaluation, plus 19.2655 ms in batch/sampling work. A replacement
must beat 7.2623 ms/evaluation before integration. To earn a useful 20 ms
end-to-end prize, the complete decoder path must be at most 202.611 ms on that
old fixture, then pass the frozen dense 15/30-second suite.

## 8. Placement and Measurement

Every promoted receipt needs:

- exact model, compiled tree, converter, runner, OS build, and host identities;
- requested compute units plus `MLComputePlan` placement;
- CPU-only and explicit-cache negative controls;
- resident warm repetitions with load/specialization reported separately;
- `os_signpost` intervals around each token and the complete decoder loop;
- Instruments or equivalent runtime corroboration;
- quiet-host admission, nominal thermal state, and raw repetitions;
- final-text and token-sequence hashes for every fixture.

Use `powermetrics` only where privilege is explicitly available and record its
exact supported sampler list. A guessed `ane_power` command is not a portable
contract across macOS versions.

## 9. Failure Modes

| Symptom | Likely cause | Falsifier |
| --- | --- | --- |
| First utterance passes, second diverges | State reused across utterances | Allocate a fresh state per utterance |
| Parallel decode becomes nondeterministic | Same state used concurrently | Serialize calls or use independent states |
| Correct output but no speedup | Full-state traffic or host dispatch dominates | Cache-length sweep and explicit-input control |
| Good body latency, bad full model | Vocabulary head or embedding gather fallback | Body/head split compute plans |
| Repeated or early EOT | Cache-position or suppression-policy mismatch | Forced known prefix and stepwise token comparison |
| Package small but memory high | Runtime weights are duplicated/respecialized | Compiled inventory and resident-memory trace |
| ANE requested but CPU busy | Silent partition or unsupported op | Compute plan plus CPU-only negative control |

Do not encode undocumented alignment multiples, a universal 2.3 ms dispatch
tax, or batch-size-one-only behavior as requirements. The raw report stated
these as laws without adequate primary evidence. Treat them as probe axes.

## 10. Decision Rule

Stateful Core ML is worth pursuing as a narrow decoder falsifier because a
20-30 ms prize may decide a marginal D750 result. It is not the primary path to
the sub-500 ms target: the stock encoder alone is approximately 622 ms and the
state API cannot delete that work.

Promote only if the full resident decoder is faster than CPU, token-exact under
the declared greedy contract, stable across independent states, and has no
material placement regression. Otherwise keep the CPU decoder and spend the
engineering budget on the encoder.

## Sources

1. [Core ML Tools: stateful models](https://apple.github.io/coremltools/docs-guides/source/stateful-models.html)
2. [Core ML Tools: multifunction models](https://apple.github.io/coremltools/docs-guides/source/multifunction-models.html)
3. [WhisperKit paper](https://arxiv.org/html/2507.10860v1)
4. [Apple Core ML `MLState`](https://developer.apple.com/documentation/coreml/mlstate)
5. [Core ML Tools source and tests](https://github.com/apple/coremltools)
6. [WhisperKitTools](https://github.com/argmaxinc/WhisperKitTools)
