# Core ML Whisper Encoder Compression and ANE Graph Surgery

> **Collective Library edition.** This is the complete technical report. Private filesystem paths, internal run identifiers, campaign-control notes, and repository navigation were removed. Technical claims, code, measurements, evidence labels, citations, corrections, and falsification criteria are preserved.

How to test low-bit weight storage and graph rewrites on a fixed-shape Whisper
encoder without confusing a smaller package, a compiler preference, or a fast
toy with a faster accurate model on an M1.

**Provenance.** Ingested 2026-08-08 from a Gemini Deep Research Max run:

```text
  core-ml-whisper-encoder-compression-and-ane-graph-surgery/
  2026-08-09T06-15-24-438Z/raw-report.md

SHA-256 77d6b4276a03989cba8920f9e06ed356c892b1d0a900bc5484c8371456700a7a
```

The Core ML Tools 9 API surface was re-checked against Apple's documentation
via Context7. Sections marked **Implementation evidence** are firsthand measurements and
take precedence over the external draft where they conflict.

## The boundary that matters

The deployed large-v3-turbo hybrid has three materially different costs:

```text
fixed 3000-frame log-mel input
        |
        v
Core ML encoder  ---- [1, 1500, 1280] ----> whisper.cpp decoder
        |                                      |
  weight traffic,                        autoregressive token
  attention, placement                   loop and KV cache
```

Compressing the Core ML package changes encoder weight storage. It does not
remove decoder tokens, audio preprocessing, VAD, or process startup. The current
campaign therefore requires an encoder below roughly 237 ms if the measured
CPU decoder and remaining compute stay around 263 ms. A package-size win is useful only if the
same compiled graph crosses that mechanical latency threshold.

> **Implementation evidence.** The fastest sealed M1 baseline is a 622.045 ms encoder and
> 884.784 ms warmed compute total on a Macmini9,1. The encoder alone is already
> 122.045 ms over the entire 500 ms goal. Residency and launch-overhead work
> cannot close that gap.

## What Core ML Tools actually provides

`coremltools.optimize.coreml.palettize_weights` rewrites constant weights in an
`mlprogram` into lookup-table storage. It is post-training and data-free: it
does not calibrate activations or establish ASR quality.

```python
import coremltools as ct
import coremltools.optimize as cto

source = ct.models.MLModel("encoder.mlpackage", skip_model_load=True)
op_config = cto.OpPalettizerConfig(
    mode="uniform",
    nbits=4,
    granularity="per_grouped_channel",
    group_size=16,
    weight_threshold=2048,
)
config = cto.OptimizationConfig(
    global_config=None,
    op_type_configs={"linear": op_config, "conv": op_config},
)
candidate = cto.palettize_weights(source, config)
candidate.save("encoder-lut4.mlpackage")
```

Core ML Tools 9 documents these relevant choices:

| Control | Documented behavior | Experimental consequence |
| --- | --- | --- |
| `mode` | `kmeans`, `uniform`, `unique`, or `custom` | Mode is part of artifact identity; never change it inside an A/B. |
| `nbits` | 1, 2, 3, 4, 6, or 8 for supported modes | Four bits means a 16-entry palette, not integer matrix math. |
| `granularity` | `per_tensor` or `per_grouped_channel` | More palettes can improve fidelity but may reduce runtime performance. |
| `group_size` | Used only for grouped-channel palettes | Group 16 is a hypothesis imported from a different graph, not a Whisper default. |
| `weight_threshold` | Defaults to 2048 elements | Small constants remain uncompressed; record coverage and bytes. |
| `cluster_dim` | Values above one enable vector palettes | Treat as a separate, newer-runtime experiment. |

Grouped-channel and vector features require the iOS 18/macOS 15 operation set.
A source `mlprogram` converted for macOS 14 cannot be assumed to accept them.
Export the experimental baseline and candidate from the same source model with
the same macOS 15 target before comparing them.

Palettization stores weights compactly and expands them through
`constexpr_lut_to_dense`. It does **not** prove that prediction executes four-bit
math, that the Neural Engine reads four times fewer bytes, or that latency falls
with package size. Those are physical questions.

The op-type filter is deliberate. Core ML Tools follows each constant to its
consumer when choosing a compression config. Restricting the candidate to
`linear` and `conv` leaves the positional embedding and other add-path
constants alone. A global config would risk quantizing the 1500x1280 positional
table for little size benefit and a disproportionate timestamp/quality risk.

## The useful Plateau transfer—and its limit

> **Implementation evidence from the sibling Plateau repository.** On its fixed-shape
> diffusion-transformer graph, `uniform` LUT4 with grouped-channel palettes and
> group size 16 compressed a four-block weight blob from about 1,745 MB to
> 436 MB while preserving Neural Engine preference for every material linear.
> A different representation—blocked INT4 from
> `linear_quantize_weights`, block size 32—moved all 28 linears to CPU even
> though the top-line report still said 81.9% Neural Engine. On an M1 Mini,
> the LUT4 arm reported 99.74% preferred Neural Engine placement.

The transferable lesson is about **representation**, not the word “grouped”:

- grouped-channel **palette** compression was the positive arm;
- grouped/blockwise **linear quantization** was the negative arm;
- top-line device percentages hid the material CPU fallback;
- the result used another architecture, tensor geometry, and in some legs
  random weights, so it says nothing yet about Whisper quality or latency.

Plateau also found only 1.03x aggregate throughput from two concurrent ANE
processes. Parallel encoders are therefore a poor primary latency bet, although
that result remains topology-specific.

## Claims the raw research did not establish

The external report contained several strong prescriptions that are not safe
to promote as facts:

- **“Every linear must become a 1x1 convolution.”** False as a universal rule.
  Layout-sensitive rewrites can help some graphs, but the existing Whisper
  package already runs and produces strong ANE power evidence. Inspect the
  exact compiled plan before rewriting trained math.
- **“A two-gigabyte cliff requires chunking this encoder.”** Plateau measured a
  cliff on another graph. The deployed Whisper encoder is already below that
  boundary. Chunking adds dispatches and intermediate transfers and needs its
  own A/B.
- **“M1 lacks blocked-quantization kernels while M3/M4 accelerate them.”** The
  public Core ML API documents operation-set availability, not this per-chip
  scheduling guarantee. Keep the Plateau result device/build-scoped.
- **“The ANE has a four-megabyte shared L2 cache.”** The supplied sources did
  not establish a public architectural contract. Do not derive Whisper chunk
  sizes from it.
- **“Change LayerNorm epsilon, masks, or softplus, then calibrate.”** Do not
  change trained semantics without a reproduced numerical failure on the exact
  Whisper graph. None of those edits is required for weight-only palettization.
- **“Quantization cannot reach the target.”** Also unproven. The current
  encoder needs a 2.62x latency improvement to fit the residual budget. A
  fourfold storage reduction is at least large enough to falsify before paying
  for architectural distillation.

> **Local graph inventory.** The current compiled Whisper encoder contains 192
> `linear`, 64 `matmul`, 32 `softmax`, and two `conv` operations, with no fused
> SDPA operation. Its weight blob is 1,273,971,776 bytes—well below Plateau's
> separate 2^31-byte cliff. The first LUT4 probe should therefore expect exactly
> 194 LUT operations for the linear/conv weights, zero LUTs for the positional
> embedding, and no attention or chunking rewrite.

## Cheapest-falsifier ladder

Run the ladder in order. A failure stops the arm; it does not authorize a more
complex rewrite.

### 1. Package physics

Build one same-source macOS 15 pair:

- FP16 baseline;
- LUT4 `uniform`, grouped-channel, group 16, threshold 2048.

Record source tree digest, Core ML Tools version, full config, package and
compiled-tree digests, bytes, operation histogram, number of
`constexpr_lut_to_dense` operations, and palette cardinality. Kill if conversion
or compilation fails, the ABI changes, the candidate does not contain exactly
194 linear/conv LUT operations, any positional-embedding constant is
palettized, or the selected weight bytes are not at least 3.5x smaller.

Do not compile this candidate through the production `compile_package` helper.
That helper correctly binds only Float16 and Float32 storage and rejects an
unknown compiled `storagePrecision`. Use raw `xcrun coremlc compile`, label the
receipt `experimental_non_certified`, archive the compiled metadata verbatim,
and record storage format (`lut4`) separately from compute precision (`fp16`).

### 2. Static placement

Load `MLComputePlan` from the compiled `.mlmodelc` under `CPU_AND_NE`. Apple's
API reports anticipated device usage and estimated cost per operation. Inspect
material `linear`, `matmul`, convolution, and attention operations individually;
do not accept an aggregate percentage.

Kill if any material candidate operation moves from a baseline Neural Engine
preference to CPU/GPU. A static plan is compiler intent, not runtime proof.

### 3. Compiled numerical parity

Run the exact compiled bundles through the same three frozen encoder probes:

- real speech fixture;
- processor-produced silence/padding;
- deterministic stress tensor.

Record max and mean absolute error, cosine similarity, and SNR per probe. Do not
borrow the production FP16 gate for LUT4 by relabeling precision. The quantized
arm needs a new experimental policy plus end-to-end transcript gates.

### 4. End-to-end quality

Use the same decoder bytes and greedy settings for both encoders. Require:

- no text regression on the public smoke and frozen pathology fixtures;
- no WER regression beyond a preregistered bound on held-out speech;
- timestamp and no-speech checks, because those can fail before normalized text;
- forbidden-log checks proving Core ML did not silently fall back.

### 5. Physical M1 latency

Run warmed, paired, rotating A/B trials on a quiet M1. Pair `MLComputePlan` with
runtime rails or an Instruments trace. Measure encoder, decoder, preprocessing,
and total separately for both approximately 15-second and 30-second clips.

The campaign go gate is mechanical:

```text
candidate encoder median < 237 ms
candidate warmed compute p50 < 500 ms
candidate warmed compute p95 < 550 ms
quality gates all pass
```

If LUT4 preserves quality and placement but misses the latency gate, its bytes
may still be useful in combination with a trained smaller encoder. It does not
win this campaign by itself.

## What to try only after LUT4

| Candidate | Why it might work | Primary risk | First kill check |
| --- | --- | --- | --- |
| Per-tensor LUT4 | Fewer palettes and simpler expansion | Worse weight fidelity | Frozen-probe and transcript parity |
| Group-16 `kmeans` LUT4 | Better fit than uniform | Slow conversion; no runtime win | Same placement and bytes, then parity |
| INT8 per-channel | Conservative compression | Only about 2x bytes, likely insufficient alone | Encoder must beat FP16 materially |
| Mixed FP16/LUT4 by op | Protects outlier-sensitive weights | Branches or exclusions erase bandwidth win | Package coverage and plan first |
| Explicit attention rewrite | Can bypass a reproduced fused-op defect | Semantic drift and graph explosion | One-layer numerical probe |
| Trained encoder compression | Can remove compute, not just bytes | Data, training, and generalization cost | Teacher-student null arm |

Do not build an automatic search over this matrix first. One exact artifact,
one exact M1, and one hard kill gate provide more information than a fleet of
unqualified package-size measurements.

## References

1. [Core ML Tools palettization API](https://apple.github.io/coremltools/source/coremltools.optimize.coreml.palettization.html)
2. [Core ML Tools grouped-channel palettization example](https://github.com/apple/coremltools/blob/main/docs-guides/source/opt-stable-diffusion.md)
3. [Core ML Tools ML Program utilities](https://apple.github.io/coremltools/source/mlmodel-utilities.html)
4. [Core ML Tools `MLComputePlan`](https://apple.github.io/coremltools/source/coremltools.models.html)
5. [Distil-Whisper paper](https://arxiv.org/abs/2311.00430)
