Stateful Core ML Whisper Decoders on Apple Silicon
Collective Library edition. This is the complete technical report. Private filesystem paths, internal run identifiers, campaign-control notes, and repository navigation were removed. Technical claims, code, measurements, evidence labels, citations, corrections, and falsification criteria are preserved.
August 9, 2026
This guide separates the documented Core ML state API from attractive but unproved claims about Neural Engine residency and base-M1 speed. It is a field guide for falsifying an internal-cache Whisper decoder, not evidence that one is faster than the existing CPU decoder.
See also the August 2026 external-research refresh, which preserves public implementation recipes while annotating them with this guide's local placement falsifier.
The raw report was corrected during ingest against Core ML Tools 9, the local macOS 15 SDK headers, the public WhisperKit paper, and pinned public WhisperKitTools source. Claims about cache traffic, operation placement, and M1 latency remain hypotheses until measured.
1. Evidence Boundary
Four statements are documented:
- Stateful
mlprogrammodels are supported from iOS 18 and macOS 15. ct.StateTypewraps a tensor state during conversion.- A runtime model can allocate an
MLStateand pass it to prediction. - Core ML Tools can represent state mutation with
read_stateandcoreml_update_state; graphs may containslice_update.
None of those statements proves that a Whisper-shaped cache stays on the ANE, that a one-slot update avoids full-state traffic, or that the resulting token loop beats whisper.cpp on a base M1.
The WhisperKit paper reports a real component result: the authors changed the large-v3-turbo decoder from cache-as-input tensors to stateful Core ML and reported a forward-pass reduction from 8.4 to 4.6 ms on an M3 ANE, or 45%. The paper does not provide the token count, timing distribution, base-M1 result, compiled package, or enough source identity to reproduce that number.
Local source correction: the pinned WhisperKitTools checkout at commit
84f77a83c8f530022ae55fbb1a64b3351ef63c7a contains no StateType or
MLState. Its decoder accepts full self-key/value caches as ordinary inputs,
returns one-token cache updates, and accepts encoder embeddings on every token.
Its pipeline defaults the decoder to cpuAndGPU. Therefore the paper is prior
art, but this public checkout is not an executable MLState reference for the
paper's 4.6 ms result.
2. Exact Runtime Contract
The macOS 15 SDK describes MLState as a handle to model-owned buffers. The
important lifecycle rules are:
newState()allocates the model's declared state buffers and initializes them to zero.prediction(from:using:)and its asynchronous equivalent consume a state handle.- Predictions sharing one state must be serialized. Reading or writing a state while a prediction uses it is undefined behavior.
withMultiArrayexposes mutable initialization/access, but the buffer address may differ across closures. Never retain a pointer beyond the closure.- There is no semantic reset operation. Allocate a fresh zeroed state for each independent utterance or decode branch.
One state per independent request is the simple safe rule. Parallel token calls on one state are invalid; parallel requests require independent states.
let configuration = MLModelConfiguration()
configuration.computeUnits = .cpuAndNeuralEngine
let model = try MLModel(contentsOf: modelURL, configuration: configuration)
let state = model.makeState()
// Serialize every prediction that shares `state`.
let output = try await model.prediction(from: input, using: state)
Requested .cpuAndNeuralEngine is a scheduler constraint, not placement
evidence. Preserve a .cpuOnly negative control and inspect the compiled
compute plan.
3. Conversion Contract
Core ML Tools 9 requires StateType.wrapped_type to be a TensorType. The
wrapped tensor cannot carry its own name or default value. The source state
name must match the traced model buffer name.
import coremltools as ct
import numpy as np
states = [
ct.StateType(
wrapped_type=ct.TensorType(
shape=(4, 20, 448, 64),
dtype=np.float16,
),
name="self_k",
),
]
model = ct.convert(
traced_model,
inputs=inputs,
states=states,
convert_to="mlprogram",
minimum_deployment_target=ct.target.macOS15,
)
For the pinned stack, treat state tensors as fixed-shape FP16 buffers. Do not generalize the raw report's statements that every dynamic input or every batch larger than one necessarily falls back; those claims were not established by the cited API. Falsify the exact shape needed by the decoder.
State update is a question, not an answer
The expected MIL shape is conceptually:
read_state -> slice_update(active position) -> coreml_update_state
That graph proves semantics, not traffic. A compiler may lower a one-slot update efficiently, materialize a larger tensor, or partition adjacent ops. Measure latency while sweeping cache length and update position. If latency scales with the full state size, the mechanism has failed even if outputs are correct.
4. Whisper Cache Arithmetic
For large-v3-turbo, the decoder has four layers, 20 heads, head dimension 64, and width 1280. With a 448-token self-cache and FP16 K/V tensors:
4 layers * 2 K/V * 20 heads * 448 positions * 64 values * 2 bytes
= 9,175,040 bytes
= 8.75 MiB
Cross-attention K/V for the full 1500 encoder positions is:
4 layers * 2 K/V * 20 heads * 1500 positions * 64 values * 2 bytes
= 30,720,000 bytes
= 29.30 MiB
The combined turbo cache is 38.05 MiB. Distil-large-v3.5 has two decoder layers, so its corresponding cache is half that size. This arithmetic does not show how many bytes the runtime moves per token.
The current turbo FP16 decoder weights are 343,818,240 bytes. At the base M1's approximate 68 GB/s memory bandwidth, one ideal stream of those weights already costs about 5.06 ms before cache reads, compute, dispatch, state update, or output handling. A 4.6 ms M3 result therefore cannot be projected directly to base M1.
5. Graph Partitioning
There are three useful experimental boundaries:
| Boundary | Question answered | What it does not prove |
|---|---|---|
| State-only update | Does update cost scale with cache size? | Decoder placement or quality |
| One decoder layer | Do attention and state mutation lower together? | Full head and package behavior |
| Full decoder | Does the product token loop win and remain exact? | Encoder or end-to-end target |
A multifunction package can combine prefill and decode functions and
deduplicate byte-identical constants. Core ML Tools documents this capability;
it does not guarantee shared runtime residency or that two functions can mutate
and consume one state exactly as desired. Prove the actual state handoff.
Separate Turbo and Distil artifacts are simplest. They have different decoder depths and learned weights. A single dynamic supergraph adds a branch and does not create useful learned-weight sharing.
Prefill and cross-attention
The attractive architecture is:
encoder output
-> prefill/setup function
-> cross K/V state
-> prompt self K/V state
-> serialized one-token decode calls
-> updated self K/V state
-> next-token decision
This is not yet a documented multifunction-state recipe. Test whether the same state object can be initialized through one function and consumed through another before designing the full decoder around it.
Whisper.cpp already owns both kv_self and read-only kv_cross. Statefulness
alone deletes no decoder math. The prize comes only if Core ML reduces actual
cache/weight traffic or executes the decoder body faster.
6. The Vocabulary Head Is a Separate Risk
The decoder body and the 51,866-way output projection have different placement and bandwidth risks. Measure these variants explicitly:
- decoder body with a small diagnostic output;
- monolithic full-vocabulary head;
- four fixed output-head shards below suspected compiler dimension cliffs;
- in-graph top-k or argmax only.
Returning only argmax is insufficient for Whisper unless the graph implements the exact suppression, no-speech, timestamp, EOT, and fallback-temperature policy owned by the product contract. For the current deterministic final-text campaign, compare the complete greedy token sequence, not merely top-1 logits on isolated steps.
Embedding gather placement and tied embedding/head storage must also be inspected. A compiled package may duplicate constants even when the source model ties them.
7. Minimal Physical Falsifier Ladder
Run these in order on the exact base M1/macOS 15 target:
7.1. State-only bandwidth probe
- Allocate exact self and cross-cache shapes.
- Compare ordinary tensor inputs with
MLState. - Sweep active cache length and update position.
- Include fresh-state, repeated-call, and two-independent-state controls.
- Record p50, p95, first-run specialization, output hashes, and compute plan.
Kill if one-slot latency scales like full-cache materialization, state reset is nondeterministic, or independent states contaminate one another.
7.2. Exact one-layer decoder
- Width 1280, 20 heads, query length one.
- Use the exact normalization, residual, self-attention, and cross-attention operations from the candidate model.
- Compare explicit cache inputs against MLState with identical weights.
- Omit the vocabulary head so placement failures are attributable.
Kill if the stateful form is not faster than the explicit form or introduces a
material CPU/GPU island under CPU_AND_NE.
7.3. Full body plus head ladder
- Scale to two layers for Distil and four for Turbo.
- Add the head variants separately.
- Run forced-prefix-to-EOT parity against the pinned CPU decoder.
- Prove cache reset across utterances and deterministic termination.
7.4. Resident end-to-end A/B
The sealed CPU decoder baseline spent 203.3455 ms on 28 decode evaluations, or 7.2623 ms/evaluation, plus 19.2655 ms in batch/sampling work. A replacement must beat 7.2623 ms/evaluation before integration. To earn a useful 20 ms end-to-end prize, the complete decoder path must be at most 202.611 ms on that old fixture, then pass the frozen dense 15/30-second suite.
8. Placement and Measurement
Every promoted receipt needs:
- exact model, compiled tree, converter, runner, OS build, and host identities;
- requested compute units plus
MLComputePlanplacement; - CPU-only and explicit-cache negative controls;
- resident warm repetitions with load/specialization reported separately;
os_signpostintervals around each token and the complete decoder loop;- Instruments or equivalent runtime corroboration;
- quiet-host admission, nominal thermal state, and raw repetitions;
- final-text and token-sequence hashes for every fixture.
Use powermetrics only where privilege is explicitly available and record its
exact supported sampler list. A guessed ane_power command is not a portable
contract across macOS versions.
9. Failure Modes
| Symptom | Likely cause | Falsifier |
|---|---|---|
| First utterance passes, second diverges | State reused across utterances | Allocate a fresh state per utterance |
| Parallel decode becomes nondeterministic | Same state used concurrently | Serialize calls or use independent states |
| Correct output but no speedup | Full-state traffic or host dispatch dominates | Cache-length sweep and explicit-input control |
| Good body latency, bad full model | Vocabulary head or embedding gather fallback | Body/head split compute plans |
| Repeated or early EOT | Cache-position or suppression-policy mismatch | Forced known prefix and stepwise token comparison |
| Package small but memory high | Runtime weights are duplicated/respecialized | Compiled inventory and resident-memory trace |
| ANE requested but CPU busy | Silent partition or unsupported op | Compute plan plus CPU-only negative control |
Do not encode undocumented alignment multiples, a universal 2.3 ms dispatch tax, or batch-size-one-only behavior as requirements. The raw report stated these as laws without adequate primary evidence. Treat them as probe axes.
10. Decision Rule
Stateful Core ML is worth pursuing as a narrow decoder falsifier because a 20-30 ms prize may decide a marginal D750 result. It is not the primary path to the sub-500 ms target: the stock encoder alone is approximately 622 ms and the state API cannot delete that work.
Promote only if the full resident decoder is faster than CPU, token-exact under the declared greedy contract, stable across independent states, and has no material placement regression. Otherwise keep the CPU decoder and spend the engineering budget on the encoder.