MissingManual

Transformer decode / latency · Collective Library

Batch-1 Transformer Decode on Apple Silicon

Collective Library edition. This is the complete technical report. Private filesystem paths, internal run identifiers, campaign-control notes, and repository navigation were removed. Technical claims, code, measurements, evidence labels, citations, corrections, and falsification criteria are preserved.

August 20, 2026

Where the per-token millisecond goes when an autoregressive decoder runs at batch 1 on an M-series chip, and which levers practitioners report actually move it. This is a domain field guide for choosing between CPU, Metal, Core ML stateful, and Neural Engine decode routes — not a verdict on any campaign arm. the sub-500 ms M1 campaign note.

The raw report was corrected during ingest against Core ML Tools 9, the local macOS 15 SDK CoreML.framework headers and Swift interface, and the ingested Apple Neural Engine architecture paper. Every correction is marked inline. External material that remains unproven is kept and labeled, because an unmeasured lever is still a lever worth ranking.

1. Evidence Boundary

Read this before quoting any number below.

Tier What it means Examples in this guide
Local measurement Sealed receipt from this repo's pinned stack on a base M1. 6.394 ms/eval CPU decoder; int8 folds to dense fp16; encoder ~100% ANE at fp16.
Corpus measurement Measured in an ingested paper under README/papers/, not by us. 0.23 ms ANE dispatch floor; 2 MB ANE working set; 260 ms idle re-wake.
Vendor documented Core ML Tools 9 source, SDK headers, or Apple docs. ct.StateType contract; MLState serialization rule; per-block-vs-per-channel guidance.
Community reported Practitioner benchmarks and issue threads from the raw report. whisper.cpp thread sweeps; Metal small-batch kernel fallback; MLX speculative-decode ratios.
Worth trying Plausible mechanism, no measurement anywhere yet. Tied-embedding CPU/GPU precision split; encoder frame stacking.

Nothing in the Community or Worth-trying tiers has been reproduced on a base M1 in this tree. They are kept because ranking unproven levers is the point of a field guide.

2. The Batch-1 Bottleneck: GEMV and the Weight-Stream Floor

2.1. Arithmetic intensity, not FLOPs

At batch 1 the decoder performs general matrix-vector multiplication. Every weight is read from DRAM to produce one token, and almost no arithmetic is done per byte read. the implementationrt puts transformer decode near 0.25 FLOP per byte. Whatever the exact figure, the consequence is fixed: decode latency scales with memory bandwidth, not with the chip's TFLOPs. Prefill is the opposite regime — it is GEMM, it hides memory latency behind parallel arithmetic, and the implementationrt cites M-series measurements putting GEMM throughput 80–85x above GEMV at all matrix sizes (Community).

The ingested ANE paper puts a hard number on where that crossover sits for one compute unit: the engine's roofline ridge point is 141 FLOP per byte. Below it a layer is bandwidth-bound against the engine's ~85 GB/s DRAM roof; above it the layer is compute-bound against a ~12 fp16 TFLOP/s roof. Batch-1 decode is three orders of magnitude below the ridge.

2.2. Achievable bandwidth is ~79–81% of advertised

the implementationrt's most portable claim: STREAM-protocol microbenchmarks on M-class silicon land at 79–81% of the advertised theoretical maximum (Community). A 153.6 GB/s part measures near 121.8 GB/s. Use the fraction, not the spec sheet, when computing a weight-streaming floor.

2.3. Local confirmation: the base-M1 turbo decoder is already at the floor

This is the one place where the implementationrt's own worked example and this repo's sealed receipt agree to the digit.

The large-v3-turbo decoder is 4 layers at d_model 1280, d_ff 5120, vocabulary 51,866, positional table 448. Its fp16 weight bytes reconstruct the receipt:

per layer  = 4·d²  (self-attn qkvo)                 =  6,553,600
           + 4·d²  (cross-attn qkvo)                =  6,553,600
           + 2·d·d_ff (mlp fc1 + fc2)               = 13,107,200
                                                      -----------
                                                      26,214,400 params
x 4 layers                                          = 104,857,600
+ token embedding 51,866 · 1,280                    =  66,388,480
+ positional      448 · 1,280                       =     573,440
                                                      -----------
                                                     171,819,520 params
x 2 bytes (fp16)                                    = 343,639,040 B
+ biases and layer norms                            =     179,200 B
                                                      -----------
                                                     343,818,240 B / token

At the base M1's 68 GB/s that is a 5.056 ms/token fp16 weight-streaming floor. The campaign note states the same figure independently: "a single ideal stream of the 343,818,240-byte FP16 turbo decoder weights already costs approximately 5.06 ms at base-M1 bandwidth."

The sealed static600 receipt (static600-stock-turbo-stop-final-v1.json) runs use_gpu: false, n_threads: 4, flash_attn: true, audio_ctx: 600, and records a decode p50 of 166.260 ms over 26 decode runs.

Provenance caution on the per-eval number. "6.394 ms/eval" is a derived quotient (166.260 ÷ 26), not a figure any receipt states. The campaign note separately states 7.262 ms per evaluation for a different arm. Both belong to the 6.4–7.3 ms/eval band, and the efficiency conclusion depends on which you use:

Per-eval basis Source Efficiency vs 5.056 ms floor
6.394 ms derived: 166.260 ms p50 ÷ 26 runs 79.1%
7.262 ms stated in the campaign note 69.6%

So the pipeline sits somewhere between 70% and 79% of the physical memory-bandwidth ceiling. The upper end lands dead centre of the 79–81% STREAM band in §2.2; the lower end still leaves under 2.2 ms/token of theoretical headroom. Quote the range, not the point estimate, unless you re-derive from a named receipt.

What this forecloses. No amount of CPU thread tuning, kernel rewriting, or scheduler work recovers meaningful time. There are exactly two exits:

  1. Stream the bytes across a wider bus (Metal/GPU), or
  2. Stream fewer bytes (quantization / palettization).

Everything else competes for a 21–30% headroom that DRAM efficiency has already largely spoken for.

2.4. Where the bytes actually are

the implementationrt's tied-embedding section is abstract. Applied to this decoder it becomes the single highest-leverage number in the guide:

Weight group fp16 bytes/token Share of stream ms at 68 GB/s
4 decoder layers 209,715,200 61.0% 3.084
Tied embedding / output projection 132,776,960 38.6% 1.953
Positional + biases + norms 1,326,080 0.4% 0.020

One matrix — the vocab projection — is 38.6% of the per-token byte stream. Palettizing only that matrix to 4-bit would cut it from 1.953 ms to 0.488 ms, removing ~1.46 ms/token from the floor while leaving all attention and MLP weights at fp16. Over a 20–40 token utterance that is 29–58 ms off the stop-to-final budget from a single-tensor change.

This is local arithmetic, not a measurement. It assumes the palettized head actually streams at 4 bits on the target runtime — which §5.1 shows is generation-dependent and must be proven per chip.

2.5. CPU threads: match P-cores, do not exceed them

the implementationrt claims the whisper.cpp sweet spot is rigidly P-core count minus one — 7 threads on an M1 Pro (8P/2E), 4 on a standard M2 (Community, [cite: 4]).

Correction (internal inconsistency). the implementationrt's own two examples disagree with its own rule: an M2 has 4 P-cores, so "P-cores minus one" would be 3, not the 4 it reports. The defensible generalization is weaker and more useful: thread count should track physical P-cores, never total logical cores. Beyond the P-core count, threads contend for the same saturated memory bus and actively degrade throughput.

Local configuration. This repo's pinned base-M1 decoder runs n_threads: 4 on a 4P/4E part — i.e. exactly the P-core count, not P-minus-one. See scripts/bench_inference_stock.py and scripts/run_static600_rolling_contention_m1.py.

3. Compute-Unit Comparison

Corrected against the ingested ANE paper. the implementationrt's ANE column was wrong in three places; the corrected values are marked.

Compute unit Per-dispatch floor Achievable bandwidth Batch-1 decode verdict
CPU (NEON/AMX) negligible (in-process) ~79–81% of system peak; AMX stores capped near 13 GB/s (Community) Works. Hits the floor. Nothing left to tune.
GPU (Metal, direct) near-zero Widest bus on the SoC; ~230 GB/s on the M1 activation stream per the ANE paper's cross-comparison The bandwidth exit. Correct home for the iterative decoder.
GPU/ANE (public Core ML) ~2.3 ms per prediction call (Community, contested — see §4.1) as above Only viable if the whole token loop is one call.
ANE (direct route) 0.23 ms per evaluation (Corpus — not the implementationrt's 70–90 µs) ~85 GB/s DRAM ceiling; ~51 GB/s measured on weight streaming; ~24 GB/s on the activation path (Corpus) Loses on physics for this decoder. See §6.3.

Correction 1 — ANE dispatch floor. the implementationrt says the raw hardware dispatch floor is "roughly 70 to 90 µs per program" via private e5rt interfaces. The ingested ANE paper measures it live on an M1 at 0.23 ms per evaluation by the slope method, and 0.19 ms for a full tiny-model call, of which ~98% is dispatch overhead and ~0.13 ms is the firmware round trip. The report is off by roughly 3x in the optimistic direction.

Correction 2 — the "32 MB SRAM cliff" does not exist. the implementationrt repeatedly claims ANE throughput "crashes by 30% if any intermediate tensor exceeds 32 MB SRAM." The ANE paper measures the on-chip operand working set at 2 MB on the M1 (1 MB on the efficiency-class engine of the M9 part). The 32 MB figure in the paper is the engine control aperture — an MMIO register window — and separately appears as the size of an example [4096, 4096] weight that exceeds the bound. the implementationrt appears to have fused the two. The real threshold is 16x smaller than stated, which changes every sizing decision downstream.

Correction 3 — ANE bandwidth is not "high FP16 TFLOPS." For a weight-streaming-bound decoder the relevant ANE number is 51 GB/s measured, not its 12 TFLOP/s compute roof. 51 GB/s is below the base M1's 68 GB/s system bandwidth. This is the crux of §6.3.

4. Dispatch Overhead and ANE Duty Cycles

4.1. The public Core ML per-call tax

MLModel.prediction crosses into a daemon over XPC and routes through IOKit. the implementationrt puts the fixed dispatch overhead at ~2.3 ms per call, with practitioners reporting ~20–24 ms per-predict under cold-start or unoptimized conditions (Community).

Unresolved conflict — flag before quoting. The research brief carried an external claim of a ~0.23 ms ANE dispatch floor per evaluation. The ingested ANE paper confirms 0.23 ms — but for the direct e5rt route, not public Core ML. the implementationrt's 2.3 ms is for the public route. These are the same digits an order of magnitude apart and are trivially confused. Treat them as two distinct numbers:

  • 0.23 ms — measured floor for a direct-route ANE program dispatch (Corpus).
  • ~2.3 ms — reported fixed cost of a public MLModel.prediction round trip (Community, unverified locally).

The gap between them is the entitlement wall. The kernel driver denies user clients without com.apple.ane.iokit-user-access, held by exactly two system binaries; every application reaches the engine through the aned broker daemon. That brokering is what the extra milliseconds buy, and it is not optional for App Store distribution.

Budget arithmetic. At 2.3 ms per call, a 20–40 token greedy decode spends 46–92 ms purely on framework communication — 9–18% of a 500 ms stop-to-final budget, before a single weight is read. Chained per-token prediction calls are therefore structurally unaffordable. The only public-API escapes are (a) fusing the whole loop into one model, or (b) MLState, which removes the cache copies but not the per-call dispatch tax.

4.2. Fusion arithmetic

The ANE paper quantifies why fusion dominates: a network run as N separate dispatches pays the 0.23 ms floor N times and copies every intermediate back to the host and forward again. Fused into one program it pays the floor once. A stack of conv-relu layers fused into one program holds per-call latency flat at ~0.19 ms from 1 layer to 32, dropping the cost charged per operation from ~222 µs to ~6.3 µs.

For a 4-layer decoder this means: if any part of the token loop lives on the ANE, all of it must, in one program. A per-layer or per-op split pays the floor once per piece and loses immediately.

4.3. Power gating and idle re-wake — report substantially wrong

the implementationrt claims the ANE sleeps after ~100 ms idle and that re-waking costs 1.2 ms to 15 ms.

Correction (Corpus). The ANE paper measures the M1 engine directly:

  • The idle state is rail-off, not clock-gated. Firmware holds the engine in ANE_POWER_STATE_ALL_OFF with every domain off until a job arrives. Idle draw reads ~0 mW.
  • Sub-second idle: a modest first-call penalty of 1.2x to 1.7x, then the next call is back to steady. (the implementationrt's "1.2 ms" appears to be a misread of a 1.2x ratio.)
  • Multi-second idle: the first call after a 5-second gap costs about 260 ms — roughly 123x the steady p50 — then the very next call returns to ~2.1 ms.

So the penalty is not 15 ms; for a dictation app with multi-second gaps between utterances it is ~260 ms, over half a 500 ms budget, paid on the first call of every burst. This confirms the research brief's external ~260 ms claim and refutes the implementationrt's.

4.4. The keep-alive pattern, and its cost

Given a 260 ms re-wake, a keep-alive is worth considering. The ANE paper's own recommendation: latency-sensitive code that must answer immediately after an idle gap should keep the engine warm with a sub-second dispatch cadence or a low-cost keep-alive call.

the implementationrt's snippet schedules a 1 Hz timer. Note that 1 Hz sits exactly at the boundary of "sub-second" and the paper's threshold for full gate-off is "a few seconds," so 1 Hz is likely sufficient but is the loosest cadence that could be.

/// Keeps the Neural Engine out of ANE_POWER_STATE_ALL_OFF between utterances.
///
/// Why this exists: on M1 the first call after a ~5 s idle gap costs ~260 ms
/// (~123x steady p50) because the compute domains gate fully off. A trivial
/// periodic dispatch keeps the base domain up so the first real token does not
/// pay that cold re-wake.
///
/// Cost: the dispatch floor draws ~0.9 W while active. On battery this is a
/// real, continuous power draw for a latency insurance policy.
let keepAlive = DispatchSource.makeTimerSource(
    queue: DispatchQueue.global(qos: .utility)
)
keepAlive.schedule(deadline: .now(), repeating: .milliseconds(800))
keepAlive.setEventHandler {
    _ = try? trivialModel.prediction(from: trivialInput)
}
keepAlive.resume()

Anti-pattern warning specific to this repo. A keep-alive timer contradicts the quiet-host gates in macos-quiet-coreml-benchmark.md. Never leave one running while collecting a benchmark receipt — it inflates co-tenancy and invalidates the quiet-host precondition.

4.5. Serialization and co-tenancy

the implementationrt claims co-tenant inflation from OS audio and camera servers contesting the same queues. The mechanism is confirmed and sharper than stated (Corpus):

  • The driver keeps at most one firmware command in flight — a single-pending-queue scheduler. Two concurrent submission threads serialize, measured at 1.04x. There is no throughput to win by overlapping.
  • The aned daemon arbitrates across all clients on one physical engine by time-division on a single request queue, holding a QoS value per request and adjusting queue depth under contention. A system dispatcher's request and yours hit the same broker.
  • The paper's blunt note on decode: "A single decode stream is serial by construction… the engine is idle between the dispatch floor and the next submission."

5. Weight-Byte Reduction Levers

Since §2.3 shows the decoder is bandwidth-bound, shrinking the weights is the latency fix. But precision formats do not all reduce streamed bytes.

5.1. int8 weight-only: documented no-op for latency, and generation-dependent

the implementationrt says the ANE datapath executes in fp16 and dequantizes int8 weights immediately before computation, so int8 weight-only quantization shrinks the file and does nothing for latency.

Confirmed three ways.

Vendor documented — Core ML Tools 9's linear_quantize_weights docstring is explicit: it converts weights to constexpr_affine_dequantize / constexpr_blockwise_shift_scale ops, and "All computation at runtime uses float precision; the precision of the intermediate tensors and the compute precision of the ops are not altered."

Local measurement — requestable int8 on M1 Core ML in this tree is weight-only and folds to dense fp16, with no latency change. Provenance caveat: this is asserted in the research brief and corroborated by placement evidence in the campaign note, but the supporting LUT4-vs-blocked-INT4 placement data in Core ML Whisper encoder compression comes from the sibling Plateau repo on a diffusion-transformer graph, not a Whisper decoder. The mechanism transfers; the specific placement percentages do not.

Corpus, with an important refinement the implementationrt misses — the ANE paper:

On the M1 the int8 fold is a stored-size saving only. The weight is half the size on disk, but it is expanded to a dense fp16 constant in DRAM before the data-movement step, so a weight-streaming-bound matmul moves full-width fp16 bytes and runs at the fp16 latency, with no bandwidth gain. The int8 weight first streams as int8 on the A14 and M2 generation, where it is dispatched as int8 and dequantized at the multiplier input rather than materialized to a dense fp16 constant in DRAM.

This is a generational boundary, not a universal law. the implementationrt presents "int8 gives zero latency benefit" as a property of the ANE. It is a property of M1-class silicon. On A14/M2 and later the int8 weight stream reaches the same effective bandwidth against its stored bytes as fp16 does — i.e. it moves half the bytes and wins. Do not port an M1 int8 conclusion forward to M2+.

5.2. Block-wise int4: real, but GPU-shaped — corrected

the implementationrt recommends block-wise int4 (block size 32) as the universal lever, claiming "block size 32 for optimal ALU mapping" and up to 4x speedup.

Correction (Core ML Tools 9 documentation). Per-block quantization is recommended for the GPU, and explicitly not the recommended granularity for the Neural Engine:

The per-block weight quantization option, introduced in iOS18/macOS15, is particularly useful for 4-bit quantization. It can provide substantial runtime memory and latency gains when models run on the GPU. Conversely, for models running on the NE, the per-channel scales option is recommended.

So the implementationrt's headline recipe is correct only for the Metal route. Routing int4 block-wise weights at the ANE is running against vendor guidance.

The parameter surface is also more flexible than the implementationrt implies. OpLinearQuantizerConfig accepts granularity in {per_tensor, per_channel, per_block}, dtype in {int4, uint4, int8, uint8}, and block_size as an int, list, or tuple — with 32 as the default, not a hardware requirement. block_size 0 means "block equals the dimension size," and an op whose dimension is not divisible by the block size is silently skipped, which is a real trap: a mis-sized block yields an uncompressed model with no error.

"""Block-wise int4 for the Metal decode route.

granularity="per_block" is the GPU recommendation. For an ANE-resident graph
Core ML Tools recommends granularity="per_channel" instead — see above.

Trap: if a weight's quantized axis is not divisible by block_size, the op is
SKIPPED, not errored. Verify the op count changed before trusting the export.
"""
import coremltools as ct
import coremltools.optimize.coreml as cto

config = cto.OptimizationConfig(
    global_config=cto.OpLinearQuantizerConfig(
        mode="linear_symmetric",
        dtype="int4",
        granularity="per_block",
        block_size=32,
    )
)
compressed = cto.linear_quantize_weights(mlmodel, config=config)

Local confirmation of the direction. int4 LUT palettization on M1 Core ML in this tree does stream fewer bytes, unlike int8. Mechanics live in Core ML Whisper encoder compression.

Kept as DATA UNAVAILABLE. the implementationrt explicitly reports no empirical LUT6 vs LUT4 latency comparison on Apple Silicon. That gap is real; do not fill it with interpolation.

5.3. Tied embeddings and mixed precision (Worth trying)

Whisper ties the token embedding to the output projection. the implementationrt proposes an asymmetric split: token lookup is memory-bound but computationally trivial and suits the CPU; the LM head is a large matrix multiply and suits the GPU. Keep embedding weights on the CPU, offload a heavily reduced (int4) LM head to the GPU, and the reduced head replaces the embedding weights in place during decode without an ID-mapping conversion (Community/paper-derived, [cite: 6, 31]).

Why this section matters more than the implementationrt realized. §2.4 shows the tied matrix is 38.6% of this decoder's per-token byte stream. Halving the head's precision is worth more than any attention optimization available. The mechanism the implementationrt describes — decoupling the lookup copy from the projection copy so they can carry different precisions — is the enabling trick, because tying forces one precision on two very differently shaped operations.

Unproven here: nobody in the corpus has measured this split on an M-series part, and untying costs a graph rewrite plus a quality re-qualification of the head. Rank it high; do not report it as a result.

6. Engine Routing: CPU, Metal, Core ML, ANE

6.1. whisper.cpp: CPU vs Metal

the implementationrt: the CPU-only ARM NEON path evaluates at 6.4–7.3 ms/eval; enabling use_gpu: true migrates to Metal and practitioners report 60–70% cuts overall (Community).

The 6.4–7.3 ms/eval figure is this repo's own number, and its provenance matters — it is a base-M1 sealed receipt, not a general M-series result. Treat the implementationrt's use of it as circular, not as independent confirmation.

Local measurement of the Metal delta. On an M1 Mini the decoder batch stage measured ~361.8 ms on CPU vs ~270 ms on Metal — a ~25% reduction, not the 60–70% the implementationrt's community sources claim. The community figure is almost certainly whole-pipeline (encoder included) on larger parts. For a base-M1 decoder, budget ~25%.

That 25% is still the single largest verified structural win available, and it is exactly the "wider bus" exit predicted by §2.3.

6.2. Flash attention

flash_attn = 1 selects a custom Metal kernel that, per the implementationrt, limits intermediate KV allocations to "effectively zero memory allocations at runtime."

Softened. "Effectively zero" is marketing phrasing for "no materialized [n_head, T, T] attention matrix." Allocations are reduced, not eliminated. The claim is directionally right and the flag is worth having on. This repo's pinned configuration sets flash_attn: true unconditionally on every arm, CPU and Metal alike.

6.3. Can the decoder run on the ANE? Physics says no on M1

the implementationrt concludes no, citing the 2.3 ms XPC tax and the (incorrect) 32 MB cliff. The conclusion is right; the arithmetic is wrong and much more decisive than stated. Using the corrected Corpus numbers on this decoder:

Bound Value Per-token cost for 343,818,240 B
ANE measured weight-streaming rate 51 GB/s 6.742 ms
ANE DRAM ceiling (best case, unreachable) 85 GB/s 4.045 ms
ANE activation path (elementwise/softmax) 24 GB/s 14.326 ms
M1 system bandwidth (what CPU/GPU see) 68 GB/s 5.056 ms

A perfectly fused, single-dispatch ANE decoder streaming at the paper's measured rate would cost 6.742 + 0.23 = ~6.97 ms/token — at or worse than the 6.4–7.3 ms/eval this repo already measures on CPU, before any Core ML brokering overhead is added. The ANE is not slow here because of scheduling. It is slow because its measured weight-streaming bandwidth (51 GB/s) is below the system bandwidth (68 GB/s) that CPU and GPU already enjoy.

Two independent structural blockers stack on top:

  • Working set. The vocab projection weight is 132.8 MB in fp16 — 63x the 2 MB on-chip operand working set. It is forced onto the memory slope and tiled regardless of any other tuning.
  • Kernel-memory budget. The ANE matmul path bounds the output-channel footprint against a 64 KB budget and rejects a matmul whose output channels do not fit. A 51,866-wide projection needs 103,732 fp16 bytes of output channel — 1.58x over — so the vocab projection cannot be expressed as a single ANE matmul at all.

The corpus verdict, stated plainly. The ANE paper's Chapter 14.10 is titled "Verdict: encoders, not decoders," and measures a single-sentence encoder ~4.4x faster than the GPU at low batch with the crossover to GPU near batch 23. Its recommendation: send the prefill, encoder, and vision front end to the engine, and send the autoregressive decode to the GPU. This matches the shipping split in large-v3-turbo ANE + CPU latency on M1 and the encoder-only ANE seam in Whisper ANEForge, which explicitly leaves the decoder in GGML.

6.4. Local placement facts

Measured in this tree, and worth knowing before any routing experiment:

  • The fp16 Core ML encoder export schedules ~100% ANE — measured as "effectively 100% of estimated compute cost on the Neural Engine with no GPU cost."
  • The fp32 export schedules on the GPU instead: 2,441 operations preferred MLGPUComputeDevice and all estimated cost landed on GPU. Precision selects the compute unit; requesting .all does not. See Core ML export debug notes.
  • The qualified decode runtimes here are HF Transformers PyTorch (the only quality-admissible baseline) and pinned whisper.cpp + fp16 Core ML encoder (the latency anchor). See KNOWN_ISSUES.

Correction to a widely repeated in-repo claim. It is often stated that "no MLState decoder export exists in this tree." That is not accurate as written. ct.StateType and make_state() do appear, in a deliberately scoped falsifier:

  • scripts/export_stateful_distil_decoder_layer.py — four ct.StateType(...) declarations exporting isolated Distil layer-0 explicit-cache and MLState probes.
  • scripts/run_stateful_distil_decoder_layer.py and tests/export/test_stateful_distil_decoder_layer.py — the runner and guard.

whisper_tuner/ itself contains no state code. The accurate statement is narrower and more informative: there is no shippable whole-decoder MLState export, and the one-layer falsifier that did run is a placement KILL. Per the campaign note: "state semantics passed, but both graphs contained material CPU islands and the stateful graph added CPU-only state machinery."

That result matters for §7: locally, the failure mode of a stateful Whisper decoder was not state semantics — those worked. It was placement.

7. Stateful Decoding and MLState

The canonical repo treatment is Stateful Core ML Whisper decoders. This section only records what ingest corrected in the implementationrt.

Read §6.4 first. The one-layer Distil MLState falsifier that ran in this tree passed state semantics and failed on placement — both the explicit-cache and stateful graphs held material CPU islands, and the stateful graph added CPU-only state machinery. That is the empirical shape of the risk here, and it is exactly the risk the implementationrt's §7.2 recipe claims to solve and cannot.

7.1. Cache-as-IO vs stateful — corrected framing

the implementationrt's table labels cache-as-IO as "$O(n^2)$ compute cost per token" and MLState as "$O(1)$ relative compute."

Correction. The cost being removed is memory copying, not compute. Cache-as-IO copies the entire KV cache in and out on every token — O(n) bytes per token, O(n²) over a sequence. Attention compute against n cached positions is O(n) per token either way; MLState does not change it. Calling it a compute saving misdirects tuning effort.

the implementationrt's cited latency pair (~1.25 tok/s → ~16.26 tok/s at a 2048 context, [cite: 24]) is a large-context LLM result. A Whisper decode of 20–40 tokens against a 448-position cache is nowhere near that regime; the copy volume being eliminated is two orders of magnitude smaller. Do not transfer the 13x.

7.2. Conversion recipe — corrected against Core ML Tools 9

the implementationrt's snippet is not runnable:

# REPORT AS WRITTEN — does not work
model = ct.convert(
    traced_model,
    states=[kv_cache_state],                      # a state OBJECT, not a StateType
    compute_units=ct.ComputeUnit.CPU_AND_NE,      # does NOT force op placement
    convert_to="mlprogram",
)                                                 # missing minimum_deployment_target

Three defects:

  1. states= takes ct.StateType instances whose name matches a key of the traced model's named_buffers(). It does not take a runtime state handle. StateType.wrapped_type must be a TensorType, and that tensor's name and default_value must both be unset or conversion raises.
  2. minimum_deployment_target=ct.target.iOS18 is required. Stateful models do not exist below it.
  3. compute_units is a load-time hint, not a placement directive. Core ML Tools' own test suite documents that the runtime "segmentor" decides dispatch and can ignore the request outright — its comment notes the GPU "supports everything but segmentor just won't choose it, even if set compute_units=ct.ComputeUnit.CPU_AND_GPU." There is no public API that forces ANE residency; the implementationrt's core "forcing ANE residency during conversion" premise is false. You can only convert, then verify (§9.1).
"""Corrected: convert a KV-cache decoder to a stateful mlprogram.

Verified against coremltools 9.0. The state name must match the traced model's
named_buffers() key exactly, or conversion silently produces a non-stateful graph.
"""
import numpy as np
import coremltools as ct

mlmodel = ct.convert(
    traced_model,
    inputs=[ct.TensorType(name="input_ids", shape=(1, 1), dtype=np.int32)],
    states=[
        ct.StateType(
            wrapped_type=ct.TensorType(shape=(n_layers, 1, max_seq_len, d_model),
                                       dtype=np.float16),
            name="k_cache",   # must equal a named_buffers() key
        ),
    ],
    minimum_deployment_target=ct.target.iOS18,  # REQUIRED for states
    convert_to="mlprogram",
)
# compute_units is a hint for THIS Python handle only; it does not pin placement.

7.3. Fixed state shapes — and the enumerated-shape rules the implementationrt got wrong

the implementationrt says stateful models bake a fixed max_seq_len, that RangeDim disables ANE, and that the fix is enumerated KV lengths.

Confirmed from the SDK header. MLStateConstraint exposes exactly one bufferShape (NSArray<NSNumber *> *) and one dataType. There is no enumerated or range variant for a state buffer. State shapes are strictly fixed at conversion. Overflowing max_seq_len is undefined behavior, as the report says.

Correction — enumerated shapes apply to inputs, not to states. You cannot give a state buffer a set of KV lengths. You can only give the inputs flexible shapes and keep the state buffer at its maximum. And the input rules are narrower than the implementationrt suggests. From coremltools/converters/_converters_entry.py:

  • Below iOS18, at most one input may use EnumeratedShapes — more raises ValueError.
  • At iOS18+, multiple enumerated-shape inputs are allowed, but all of them must declare the same number of shapes, or conversion raises.

For a decoder whose token input, causal mask, and cross-attention inputs all vary with KV length, that second rule is the binding constraint and the implementationrt does not mention it.

Reading them back at runtime uses MLMultiArrayShapeConstraint.enumeratedShapes, typed NSArray<NSArray<NSNumber *> *> — so the implementationrt's shapeConstraint.enumeratedShapes[0][1].intValue indexing is valid.

7.4. Serialization — confirmed verbatim by the SDK

the implementationrt's concurrency claim is exactly right, and the SDK header states it without hedging:

The object is a handle to the state buffers. The client shall not read or write the buffers while a prediction is in-flight. Each stateful prediction that uses the same MLState must be serialized. Otherwise, if two such predictions run concurrently, the behavior is undefined.

Allocate one MLState per concurrent stream via makeState().

/// Verified against the macOS 15 SDK Swift interface.
///
/// makeState() is macOS 15+, is NOT throwing, and is the correct Swift spelling.
/// (`newState()` exists in the ObjC header but is marked unavailable on macOS,
/// iOS, tvOS, watchOS, and visionOS — do not call it. the implementationrt's
/// `try model.makeState()` adds a `try` the signature does not need.)
let state = model.makeState()

for _ in 0 ..< maxTokens {
    // prediction(from:using:) is macOS 15+ and mutates `state` in place.
    // This loop MUST stay serial: concurrent predictions sharing one MLState
    // are undefined behavior per the SDK header.
    let output = try await model.prediction(from: input, using: state)
    // ... greedy argmax, append, build next input ...
}

7.5. Reported OS bugs (Community — unverified locally)

Kept as reported, and explicitly unverified in this tree. Treat as a pre-flight checklist for a stateful export, not as known behavior.

  • iOS 18.4/18.5 VNCoreMLTransformer faults under heavy sync-task load.
  • computeUnits = .cpuOnly quietly breaking numerical decode; .all crashing at model load.
  • Greedy output drifting across days/OS sessions on identical inputs, attributed to GPU shader recompilation.
  • Symmetric int8 failing on Metal with an MPSGraphExecutable.mm assertion.

The .cpuOnly item has a documented mechanism worth knowing: Core ML Tools' test suite notes that cpuOnly is not one backend — it dispatches across BNNS and classic CPU, where "BNNS does not support sdpa" and "Classic CPU does not support int8 weight." A model can therefore fail or degrade on .cpuOnly for reasons that have nothing to do with the CPU being slow.

8. Speculative and Parallel Decoding at Batch 1

the implementationrt's finding is the useful one, and it is unusually well aligned with this repo's history.

The unified-memory problem. Speculative decoding's arithmetic assumes verifying K tokens is nearly free because the arithmetic units idle while weights stream. the implementationrt cites M-series measurements where a 2-token verification forward pass costs nearly twice a 1-token forward pass (Community, [cite: 35]) — i.e. the batched verify stays bandwidth-bound and the "free" assumption collapses. Reported Apple-specific end-to-end gains max out at 1.05x–1.6x (MLX benchmarks, Community), and for small target models the draft model's own cost can produce a net deceleration.

Metal small-batch kernel trap (Community, [cite: 5]). Optimized weight-reorder kernels are bootstrapped for the single-token case (ne[1] == 1). When a verification step passes a small multi-column array (ne[1] between 2 and 16, e.g. ne[1] <= 8), the GPU quietly falls back to a slower non-reorder kernel. The verification batch — the entire point of speculation — lands precisely in the shape range with the worst kernel selection. This is worth checking directly before attributing a speculative-decode failure to acceptance rate.

Local corroboration — two speculative arms already died here. Both were killed on preregistered measurement, independently of this report:

  • CTC draft head — Plan 09, preregistered Phase 2 KILL: "the head accepted 0/7,189 canonical draft tokens: every non-empty draft mismatched at position zero." Narration in the CTC draft note.
  • Shared-first-layer self-speculative draft — Plan 10, Phase 1 KILL: "The exactness gates passed and every accepted-depth gate failed." Note the shape of that result — correctness was never the problem; acceptance depth was.

the implementationrt's mechanism explains why a third attempt should not be cheap to justify: even at a perfect acceptance rate, a 4-layer decoder emitting 20–40 tokens has to amortize draft execution and verification-batch overhead across a burst too short to pay for either.

Two adjacent local kills belong on the same shelf:

  • Append-only decoder cache — KILL: "Appending the corresponding decoder cross-attention K/V does not make provisional decoder self-attention K/V reusable." Mechanism in Causal cached Whisper on Apple Silicon.
  • FP32 MLX decoder on M1 — KILL: "The M1 p95 is 3.60x that budget."

And one kill that is routinely over-generalized: the Alpha16 whisper.cpp F16/Q5/true-FP32 work was killed checkpoint-scoped — "the checkpoint-750 whisper.cpp runtime path is killed, not merely the compressed formats." That is a kill of that runtime path on that checkpoint, not a general verdict on GGML quantization. Do not cite it as one.

9. Debugging Playbook

9.1. Silent CPU fallback

Symptom: MLState or decoder layers evaluate at >20 ms/token; a model you believe is ANE-resident is not.

Cause: Core ML hit an unsupported op, a RangeDim dynamic shape, or an unoptimized scatter-gather state tensor, and re-segmented onto another unit. No error is raised. Since compute_units is only a hint (§7.2), this is the normal failure mode, not an exceptional one.

Verification — corrected snippet. the implementationrt's MLComputePlan example does not compile. Verified against the macOS 15 SDK Swift interface:

/// Prove which compute device each operation actually landed on.
///
/// Corrections vs the commonly circulated snippet:
///   - MLComputePlan is macOS 14.4 / iOS 17.4, NOT macOS 15 / iOS 18.
///   - `modelStructure` is an ENUM (.program / .neuralNetwork / .pipeline /
///     .unsupported). There is no `modelStructure.operations`.
///   - Operations live at program.functions["main"].block.operations.
///   - There is no `op.preferredComputeUnit`. Ask the PLAN via
///     `plan.deviceUsage(for:)`, which returns `.preferred` and `.supported`
///     as MLComputeDevice values (not compute *units*).
@available(macOS 14.4, iOS 17.4, *)
func auditPlacement(modelURL: URL, configuration: MLModelConfiguration) async throws {
    let plan = try await MLComputePlan.load(contentsOf: modelURL,
                                            configuration: configuration)

    guard case let .program(program) = plan.modelStructure,
          let main = program.functions["main"] else {
        print("Not an mlprogram — nothing to audit.")
        return
    }

    for op in main.block.operations {
        guard let usage = plan.deviceUsage(for: op) else { continue }
        // .weight is a relative cost share of the whole plan, useful for
        // ranking which misplaced op actually costs you anything.
        let cost = plan.estimatedCost(of: op)?.weight ?? 0
        print("\(op.operatorName)  preferred=\(usage.preferred)  "
              + "supported=\(usage.supported)  weight=\(cost)")
    }
}

Cross-check with the Xcode Core ML Performance Report. Per KNOWN_ISSUES, ANE residency is not claimed in this repo without MLComputePlan or Instruments evidence.

9.2. Latency floor reality check

Symptom: Kernel or threading work produces no measurable improvement.

Diagnosis: Compute the weight-stream floor before optimizing anything.

"""Is this decoder bandwidth-bound, or is there real headroom?

If measured/floor lands near the 1.25x mark, the implementation is already at
the 79-81% STREAM efficiency of the hardware and no kernel work will help.
Reduce bytes or widen the bus instead.
"""
BYTES_PER_TOKEN = 343_818_240      # fp16 large-v3-turbo decoder, incl. tied head
SYSTEM_BANDWIDTH_BPS = 68e9        # base M1 LPDDR4X

floor_ms = BYTES_PER_TOKEN / SYSTEM_BANDWIDTH_BPS * 1e3   # 5.056 ms

# Derive per-eval from a named receipt; do not hardcode a remembered number.
# static600-stock-turbo-stop-final-v1.json: decode_ms.p50 / decode_runs
per_eval_ms = 166.260 / 26                                 # 6.394 ms
print(f"efficiency = {floor_ms / per_eval_ms:.1%}")        # 79.1% -> at the ceiling

9.3. Quantization that did nothing

Symptom: An int8 export is half the size on disk and exactly as slow.

Cause: Expected on M1. The weight expands to a dense fp16 constant in DRAM before the data-movement step (§5.1). Nothing about this is a bug.

Fix: Use 4-bit palettization/quantization, which actually reduces streamed bytes. Verify the compressed ops exist before benchmarking — a block size that does not divide the quantized axis causes the op to be skipped silently.

9.4. Jitter correlated with audio capture

Symptom: whisper.cpp eval times fluctuate; AVFoundation drops frames.

Cause: Decode threads contend with the OS audio server for P-cores (§2.5), or ANE work queues behind another client at the shared aned broker (§4.5).

Verification: Confirm capture delegates fire on schedule; pin inference to .userInitiated or isolate behind an actor; re-run under the quiet-host gates in macos-quiet-coreml-benchmark.md.

10. Anti-Patterns

Anti-pattern Why it fails Do this instead
Tuning CPU threads to fix decode latency Already at 79.1% of the memory ceiling; threads only add contention Reduce streamed bytes, or move to Metal
Setting threads to total logical cores E-cores contend for the same saturated bus Track physical P-core count
Chained per-token public MLModel.prediction ~2.3 ms/call × 20–40 tokens = 46–92 ms of pure framework overhead Fuse the loop into one model, or accept a non-Core-ML runtime
compute_units=CPU_AND_NE to "force" ANE residency A load hint the segmentor may ignore; no public API pins placement Convert, then verify with MLComputePlan (§9.1)
int8 weight-only for M1 latency Expands to dense fp16 in DRAM; zero bandwidth gain 4-bit palettization / block-wise int4
Porting an M1 int8 conclusion to M2+ A14/M2 stream int8 natively and do win Re-measure per silicon generation
Block-wise int4 on an ANE-resident graph Core ML Tools recommends per-channel for the NE granularity="per_channel" for ANE, per_block for GPU
Running the AR token loop on the ANE 51 GB/s < 68 GB/s system bandwidth; 63x over the 2 MB working set; vocab matmul exceeds the 64 KB output-channel budget Encoder/prefill on ANE, decode on GPU
Enumerated shapes on a state buffer MLStateConstraint has exactly one fixed bufferShape Fix the state at max_seq_len; flex the inputs
Multiple enumerated-shape inputs with differing shape counts Raises at conversion (iOS18+ rule) Give every enumerated input the same number of shapes
Concurrent predictions sharing one MLState Undefined behavior per the SDK header One MLState per stream; serialize within a stream
Keep-alive timer during a benchmark Violates the quiet-host precondition; inflates co-tenancy Keep-alive in production only, never while sealing a receipt
Assuming speculative verify of K tokens is free 2-token verify measured near 2x a 1-token verify on M-series Model the verify as bandwidth-bound; check acceptance first

11. Decision Framework

For a 4-layer, 1280-wide decoder emitting 20–40 greedy tokens on a base M1.

Route Expected effect Cost / risk Use when
CPU (current) 6.394 ms/eval, 79.1% of floor None; already qualified Baseline. Nothing left to tune.
Metal via whisper.cpp use_gpu ~25% locally measured on the decoder batch stage Kernel-selection and parity risk The one verified structural win. Try first.
4-bit on the tied head only ~1.46 ms/token off the floor (38.6% of stream) Quality re-qualification of the output head Highest byte-reduction leverage per unit of work.
Block-wise int4, whole decoder Up to ~4x byte reduction on the GPU path Quality risk across all weights; ANE wants per-channel After the head-only result is known.
Core ML stateful decoder Removes cache copies, not the per-call tax No MLState export exists here; large-context gains don't transfer to 448 positions Only alongside a fused single-call loop.
ANE decode Physics-losing on M1 (~6.97 ms/token best case) Two structural blockers on top Do not. Encoder only.
Speculative decode 1.05–1.6x elsewhere; 2 arms already dead here Draft cost unamortizable over 20–40 tokens Not without a new acceptance-rate result.
Encoder frame-rate reduction Cuts cross-attention KV by up to 75% Retraining; acoustic-fidelity risk If first-token prefill, not per-token decode, is the bottleneck.

12. Appendix: Encoder Frame-Rate Reduction (Worth trying)

If per-token decode cannot be squeezed further, reduce what the decoder must attend to. Whisper encoders emit at 50 Hz — 1500 positions for a 30-second clip.

the implementationrt describes adapting Zipformer-style frame stacking: 1D causal convolutions or multi-scale dilated convolutions compress 50 Hz to 25 Hz or 12.5 Hz (4x stacking) (Community, [cite: 44, 45]). Reported quality cost is low — research retrofitting Whisper with frame-stacking downsampling indicates acoustic fidelity is largely preserved, especially if GELU non-linearities are removed from the initial convolutional layers, which otherwise suppress fine spectral detail ([cite: 45, 46]).

Where the win lands, precisely. Downsampling 1500 → 375 positions cuts the cross-attention KV cache by 75%, which reduces cross-attention FLOPs quadratically and improves first-token prefill. It does not reduce the per-token weight stream from §2.4 — the decoder's own weights are unchanged. So this lever targets a different line item in the stop-to-final budget than everything else in this guide. See Whisper streaming stop-to-final latency for how the two compose.

Related local work on shrinking encoder work lives in Whisper full-context encoder students on ANE.

Works Cited

Primary sources verified during ingest:

  1. Core ML Tools — stateful models — ct.StateType, states=, minimum_deployment_target=ct.target.iOS18.
  2. Core ML Tools — OpLinearQuantizerConfig and linear_quantize_weights — granularity options, block_size default 32 and skip-on-indivisible behavior, per-block for GPU vs per-channel for NE, and "all computation at runtime uses float precision."
  3. macOS 15 SDK, CoreML.framework/Headers/MLState.h — the verbatim serialization requirement.
  4. macOS 15 SDK, CoreML.framework/Headers/MLStateConstraint.h — the single fixed bufferShape.
  5. macOS 15 SDK, CoreML.swiftmodule interface — makeState() availability, newState() unavailability on macOS, MLComputePlan.load(contentsOf:configuration:) at macOS 14.4, MLModelStructure as an enum, deviceUsage(for:).
  6. coremltools/converters/_converters_entry.py (9.0) — _validate_enumerated_shape_inputs single-input and equal-shape-count rules.
  7. Apple Neural Engine: Architecture, Programming, and Performance — 0.23 ms dispatch floor, 2 MB working set, 64 KB output-channel budget, 51/85/24 GB/s bandwidth figures, rail-off idle and 260 ms re-wake, 1.04x submission serialization, "encoders, not decoders."

Secondary and community sources from the raw report are preserved in the raw export at the provenance path above, with [cite: N] markers retained inline where a specific claim depends on one. They were not independently verified during ingest and are labeled Community throughout.