Core ML Whisper Encoder Compression and ANE Graph Surgery
Collective Library edition. This is the complete technical report. Private filesystem paths, internal run identifiers, campaign-control notes, and repository navigation were removed. Technical claims, code, measurements, evidence labels, citations, corrections, and falsification criteria are preserved.
How to test low-bit weight storage and graph rewrites on a fixed-shape Whisper encoder without confusing a smaller package, a compiler preference, or a fast toy with a faster accurate model on an M1.
Provenance. Ingested 2026-08-08 from a Gemini Deep Research Max run:
core-ml-whisper-encoder-compression-and-ane-graph-surgery/
2026-08-09T06-15-24-438Z/raw-report.md
SHA-256 77d6b4276a03989cba8920f9e06ed356c892b1d0a900bc5484c8371456700a7a
The Core ML Tools 9 API surface was re-checked against Apple's documentation via Context7. Sections marked Implementation evidence are firsthand measurements and take precedence over the external draft where they conflict.
The boundary that matters
The deployed large-v3-turbo hybrid has three materially different costs:
fixed 3000-frame log-mel input
|
v
Core ML encoder ---- [1, 1500, 1280] ----> whisper.cpp decoder
| |
weight traffic, autoregressive token
attention, placement loop and KV cache
Compressing the Core ML package changes encoder weight storage. It does not remove decoder tokens, audio preprocessing, VAD, or process startup. The current campaign therefore requires an encoder below roughly 237 ms if the measured CPU decoder and remaining compute stay around 263 ms. A package-size win is useful only if the same compiled graph crosses that mechanical latency threshold.
Implementation evidence. The fastest sealed M1 baseline is a 622.045 ms encoder and 884.784 ms warmed compute total on a Macmini9,1. The encoder alone is already 122.045 ms over the entire 500 ms goal. Residency and launch-overhead work cannot close that gap.
What Core ML Tools actually provides
coremltools.optimize.coreml.palettize_weights rewrites constant weights in an
mlprogram into lookup-table storage. It is post-training and data-free: it
does not calibrate activations or establish ASR quality.
import coremltools as ct
import coremltools.optimize as cto
source = ct.models.MLModel("encoder.mlpackage", skip_model_load=True)
op_config = cto.OpPalettizerConfig(
mode="uniform",
nbits=4,
granularity="per_grouped_channel",
group_size=16,
weight_threshold=2048,
)
config = cto.OptimizationConfig(
global_config=None,
op_type_configs={"linear": op_config, "conv": op_config},
)
candidate = cto.palettize_weights(source, config)
candidate.save("encoder-lut4.mlpackage")
Core ML Tools 9 documents these relevant choices:
| Control | Documented behavior | Experimental consequence |
|---|---|---|
mode |
kmeans, uniform, unique, or custom |
Mode is part of artifact identity; never change it inside an A/B. |
nbits |
1, 2, 3, 4, 6, or 8 for supported modes | Four bits means a 16-entry palette, not integer matrix math. |
granularity |
per_tensor or per_grouped_channel |
More palettes can improve fidelity but may reduce runtime performance. |
group_size |
Used only for grouped-channel palettes | Group 16 is a hypothesis imported from a different graph, not a Whisper default. |
weight_threshold |
Defaults to 2048 elements | Small constants remain uncompressed; record coverage and bytes. |
cluster_dim |
Values above one enable vector palettes | Treat as a separate, newer-runtime experiment. |
Grouped-channel and vector features require the iOS 18/macOS 15 operation set.
A source mlprogram converted for macOS 14 cannot be assumed to accept them.
Export the experimental baseline and candidate from the same source model with
the same macOS 15 target before comparing them.
Palettization stores weights compactly and expands them through
constexpr_lut_to_dense. It does not prove that prediction executes four-bit
math, that the Neural Engine reads four times fewer bytes, or that latency falls
with package size. Those are physical questions.
The op-type filter is deliberate. Core ML Tools follows each constant to its
consumer when choosing a compression config. Restricting the candidate to
linear and conv leaves the positional embedding and other add-path
constants alone. A global config would risk quantizing the 1500x1280 positional
table for little size benefit and a disproportionate timestamp/quality risk.
The useful Plateau transfer—and its limit
Implementation evidence from the sibling Plateau repository. On its fixed-shape diffusion-transformer graph,
uniformLUT4 with grouped-channel palettes and group size 16 compressed a four-block weight blob from about 1,745 MB to 436 MB while preserving Neural Engine preference for every material linear. A different representation—blocked INT4 fromlinear_quantize_weights, block size 32—moved all 28 linears to CPU even though the top-line report still said 81.9% Neural Engine. On an M1 Mini, the LUT4 arm reported 99.74% preferred Neural Engine placement.
The transferable lesson is about representation, not the word “grouped”:
- grouped-channel palette compression was the positive arm;
- grouped/blockwise linear quantization was the negative arm;
- top-line device percentages hid the material CPU fallback;
- the result used another architecture, tensor geometry, and in some legs random weights, so it says nothing yet about Whisper quality or latency.
Plateau also found only 1.03x aggregate throughput from two concurrent ANE processes. Parallel encoders are therefore a poor primary latency bet, although that result remains topology-specific.
Claims the raw research did not establish
The external report contained several strong prescriptions that are not safe to promote as facts:
- “Every linear must become a 1x1 convolution.” False as a universal rule. Layout-sensitive rewrites can help some graphs, but the existing Whisper package already runs and produces strong ANE power evidence. Inspect the exact compiled plan before rewriting trained math.
- “A two-gigabyte cliff requires chunking this encoder.” Plateau measured a cliff on another graph. The deployed Whisper encoder is already below that boundary. Chunking adds dispatches and intermediate transfers and needs its own A/B.
- “M1 lacks blocked-quantization kernels while M3/M4 accelerate them.” The public Core ML API documents operation-set availability, not this per-chip scheduling guarantee. Keep the Plateau result device/build-scoped.
- “The ANE has a four-megabyte shared L2 cache.” The supplied sources did not establish a public architectural contract. Do not derive Whisper chunk sizes from it.
- “Change LayerNorm epsilon, masks, or softplus, then calibrate.” Do not change trained semantics without a reproduced numerical failure on the exact Whisper graph. None of those edits is required for weight-only palettization.
- “Quantization cannot reach the target.” Also unproven. The current encoder needs a 2.62x latency improvement to fit the residual budget. A fourfold storage reduction is at least large enough to falsify before paying for architectural distillation.
Local graph inventory. The current compiled Whisper encoder contains 192
linear, 64matmul, 32softmax, and twoconvoperations, with no fused SDPA operation. Its weight blob is 1,273,971,776 bytes—well below Plateau's separate 2^31-byte cliff. The first LUT4 probe should therefore expect exactly 194 LUT operations for the linear/conv weights, zero LUTs for the positional embedding, and no attention or chunking rewrite.
Cheapest-falsifier ladder
Run the ladder in order. A failure stops the arm; it does not authorize a more complex rewrite.
1. Package physics
Build one same-source macOS 15 pair:
- FP16 baseline;
- LUT4
uniform, grouped-channel, group 16, threshold 2048.
Record source tree digest, Core ML Tools version, full config, package and
compiled-tree digests, bytes, operation histogram, number of
constexpr_lut_to_dense operations, and palette cardinality. Kill if conversion
or compilation fails, the ABI changes, the candidate does not contain exactly
194 linear/conv LUT operations, any positional-embedding constant is
palettized, or the selected weight bytes are not at least 3.5x smaller.
Do not compile this candidate through the production compile_package helper.
That helper correctly binds only Float16 and Float32 storage and rejects an
unknown compiled storagePrecision. Use raw xcrun coremlc compile, label the
receipt experimental_non_certified, archive the compiled metadata verbatim,
and record storage format (lut4) separately from compute precision (fp16).
2. Static placement
Load MLComputePlan from the compiled .mlmodelc under CPU_AND_NE. Apple's
API reports anticipated device usage and estimated cost per operation. Inspect
material linear, matmul, convolution, and attention operations individually;
do not accept an aggregate percentage.
Kill if any material candidate operation moves from a baseline Neural Engine preference to CPU/GPU. A static plan is compiler intent, not runtime proof.
3. Compiled numerical parity
Run the exact compiled bundles through the same three frozen encoder probes:
- real speech fixture;
- processor-produced silence/padding;
- deterministic stress tensor.
Record max and mean absolute error, cosine similarity, and SNR per probe. Do not borrow the production FP16 gate for LUT4 by relabeling precision. The quantized arm needs a new experimental policy plus end-to-end transcript gates.
4. End-to-end quality
Use the same decoder bytes and greedy settings for both encoders. Require:
- no text regression on the public smoke and frozen pathology fixtures;
- no WER regression beyond a preregistered bound on held-out speech;
- timestamp and no-speech checks, because those can fail before normalized text;
- forbidden-log checks proving Core ML did not silently fall back.
5. Physical M1 latency
Run warmed, paired, rotating A/B trials on a quiet M1. Pair MLComputePlan with
runtime rails or an Instruments trace. Measure encoder, decoder, preprocessing,
and total separately for both approximately 15-second and 30-second clips.
The campaign go gate is mechanical:
candidate encoder median < 237 ms
candidate warmed compute p50 < 500 ms
candidate warmed compute p95 < 550 ms
quality gates all pass
If LUT4 preserves quality and placement but misses the latency gate, its bytes may still be useful in combination with a trained smaller encoder. It does not win this campaign by itself.
What to try only after LUT4
| Candidate | Why it might work | Primary risk | First kill check |
|---|---|---|---|
| Per-tensor LUT4 | Fewer palettes and simpler expansion | Worse weight fidelity | Frozen-probe and transcript parity |
Group-16 kmeans LUT4 |
Better fit than uniform | Slow conversion; no runtime win | Same placement and bytes, then parity |
| INT8 per-channel | Conservative compression | Only about 2x bytes, likely insufficient alone | Encoder must beat FP16 materially |
| Mixed FP16/LUT4 by op | Protects outlier-sensitive weights | Branches or exclusions erase bandwidth win | Package coverage and plan first |
| Explicit attention rewrite | Can bypass a reproduced fused-op defect | Semantic drift and graph explosion | One-layer numerical probe |
| Trained encoder compression | Can remove compute, not just bytes | Data, training, and generalization cost | Teacher-student null arm |
Do not build an automatic search over this matrix first. One exact artifact, one exact M1, and one hard kill gate provide more information than a fleet of unqualified package-size measurements.