Physical whisper.cpp + Core ML Encoder Verification on Apple Silicon (2026)
Collective Library edition. This is the complete technical report. Private filesystem paths, internal run identifiers, campaign-control notes, and repository navigation were removed. Technical claims, code, measurements, evidence labels, citations, corrections, and falsification criteria are preserved.
August 1, 2026
This guide defines how to prove, from a physical run, that a pinned whisper.cpp build actually executed a Core ML encoder instead of its ordinary GGML encoder. It exists because the Phase 7 release gate treats a silent ordinary-encoder run as a FAILED gate, and because a passing transcript is not evidence of which encoder produced it.
Scope: the runtime boundary only — build flags, artifact naming, the exact log contract, the headless subprocess contract, parity binding, and compute-device observation. Model conversion theory lives in the Core ML guides; the container question lives in the runtime-contract guide.
Verified against:
| Source | Revision / version | Checked |
|---|---|---|
| whisper.cpp (repo pin) | 97c56f1dc1d1100a9d859c865a20c82d22f823ed plus patch 99bf1e4a9f768745bbaad61d4d131e5360c2f240e2fa577ce64726d9c67bfcef |
2026-08-01 |
whisper.cpp master |
2ca53bb45e38748d07b310eeb36245a7157ac882 |
2026-08-01 |
| coremltools | 9.0 (released 2025-11-10; current latest on PyPI) |
2026-07-30 |
| Hugging Face Transformers | main (modeling_whisper.py) |
2026-07-30 |
| Host used for local tool checks | macOS 15.7.7 (24G720), Xcode installed | 2026-07-30 |
At both checked source revisions, whisper_state.batch is still declared
without initialization, the required Core ML failure returns before
whisper_batch_init, and teardown unconditionally calls whisper_batch_free.
Do not infer whole-file identity from that lifecycle match; line numbers below
name the pinned source unless stated otherwise.
This guide was refreshed from the Deep Research Max draft at
The draft's claim that issue #3807 documents this defect was removed: #3807 is
about unvalidated zero-dimension model parameters, not whisper_batch teardown.
The lifecycle claims retained here were checked against the pinned and current
upstream source plus the implementationsitory's repeated physical A/B.
1. Direct verdict
| Question | Verified answer |
|---|---|
| Which single log line proves the Core ML encoder loaded? | whisper_init_state: Core ML model loaded |
| Does that line also prove the GGML encoder was skipped? | Yes. The same pointer that gates the line gates whisper_encode_external(). See section 2.2. |
| Does it prove the Apple Neural Engine ran? | No. Nothing in whisper.cpp observes device placement. |
| Where do the log lines go? | stderr, always. Never stdout. |
| Can the gate lose the log lines? | Yes — whisper-cli -np/--no-prints installs a null log callback and silences them. |
| Is silent fallback possible by default? | Not from a load failure: WHISPER_COREML_ALLOW_FALLBACK defaults OFF, so init hard-fails. It is possible from a build that simply lacks WHISPER_COREML=1. |
| What does the build-level marker look like? | COREML = 1 inside the system_info: line. COREML = 0 (or an absent marker) means the binary has no Core ML code at all. |
| Which encoder input shape does the runtime hand to Core ML? | [1, n_mels, 2 * n_audio_ctx] = [1, 80, 3000] for standard models. |
| Which compute units does whisper.cpp request? | MLComputeUnitsAll, hardcoded. Not configurable without patching source. |
Is .mlpackage parity the same evidence as .mlmodelc parity? |
No. The runtime loads the compiled .mlmodelc; bind parity to those exact bytes. |
| What identifies the supported downstream runtime? | Upstream commit + exact compatibility-patch SHA-256 + built whisper-cli SHA-256. |
| May the negative control crash? | No. It must exit exactly 3 and emit both the Core ML load-failure and final CLI initialization-failure markers. |
2. The load-confirmation contract
2.1. Every Core ML line the runtime can emit
All four come from whisper_init_state(), so __func__ expands to
whisper_init_state. Each row's text is the corresponding source format string
with %s substituted — the literal portions are exact, and the only variable
part is the path:
Emitted text (after %s expansion) |
Source | Meaning |
|---|---|---|
whisper_init_state: loading Core ML model from '<path>' |
src/whisper.cpp:3443 |
Core ML code is compiled in and a path was derived. Proves nothing about success. |
whisper_init_state: first run on a device may take a while ... |
src/whisper.cpp:3444 |
Always printed, unconditionally, before the load attempt. Not a latency measurement. |
whisper_init_state: failed to load Core ML model from '<path>' |
src/whisper.cpp:3448 |
Load failed. Logged at ERROR level. |
whisper_init_state: Core ML model loaded |
src/whisper.cpp:3454 |
The load-confirmation line. Emitted only on the else branch of if (!state->ctx_coreml). |
Build-level marker, from whisper_print_system_info():
s += "COREML = " + std::to_string(whisper_has_coreml()) + " | ";
src/whisper.cpp:4334, where whisper_has_coreml() (src/whisper.cpp:4313)
returns 1 only under #ifdef WHISPER_USE_COREML. whisper-cli embeds the
whole string in one system_info: line (examples/cli/cli.cpp:1187–1188).
The line's composition, derived from those two sources rather than quoted from
a run: system_info: n_threads = <n> / <m> | then WHISPER : COREML = <0|1> | OPENVINO = <0|1> | then one segment per registered ggml backend with its
feature flags. The backend tail varies by machine and by build, and the
upstream README's example still shows a pre-backend-registry format. Match on
the substring COREML = 1. Never match the whole line.
2.2. Why the load line is sufficient proof of encoder substitution
This is the chain the gate rests on, and it is airtight in current source:
state->ctx_coremlis assigned atsrc/whisper.cpp:3446.whisper_init_state: Core ML model loadedis printed iffstate->ctx_coreml != nullptr(src/whisper.cpp:3447–3454).whisper_encode_external()returnsuse_coreml, defined aswstate.ctx_coreml != nullptr(src/whisper.cpp:1964), i.e. exactly the same predicate.- In
whisper_encode_internal(), whenwhisper_encode_external()is true the conv graph compute is skipped andwhisper_coreml_encode(...)runs instead (src/whisper.cpp:2405–2413); the encoder graph block is then also skipped (src/whisper.cpp:2421). - In
whisper_build_graph_conv(), the external branch does not even build the convolution ops — it allocates a bareembd_encinput tensor "the external encoder will write into" (src/whisper.cpp:2020–2026).
So the load marker is not a hint. It is the same boolean that removes the GGML encoder from the graph.
2.3. The assertion the Phase 7 gate must make
Four conditions, all required, from one physical run:
- exit status
0; COREML = 1present (build actually has Core ML compiled in);whisper_init_state: Core ML model loadedpresent;whisper_init_state: failed to load Core ML modelabsent.
Condition 4 matters only if a build ever sets WHISPER_COREML_ALLOW_FALLBACK,
but asserting it costs nothing and converts a misconfigured build from a silent
pass into a loud failure.
the implementationsitory encodes conditions 2–4 as
COREML_ENABLED_MARKER, COREML_LOAD_MARKER, and
COREML_FALLBACK_FAILURE_MARKER. The shared failure-marker and termination
contract lives in whisper_tuner/export/runtime_contracts.py; the physical
probe lives in whisper_tuner/export/coreml_fallback_probe.py.
2.4. The negative control
Assertions on a passing run can be satisfied by a stale log. Pair every gate
with a fail-closed canary: copy only the ggml-*.bin into a scratch
directory — deliberately leaving the .mlmodelc behind — and run the same
binary. Required outcome:
- exit exactly
3; failed to load Core ML modelpresent;error: failed to initialize whisper contextpresent;- no retry and no signal termination (
-6,134, or otherwise).
Chain, verified in source: whisper_init_state returns nullptr when Core ML
fails and fallback is not enabled (src/whisper.cpp:3449–3452) →
whisper_init_from_file_with_params returns nullptr
(src/whisper.cpp:3755–3759) → whisper-cli prints
error: failed to initialize whisper context and returns 3
(examples/cli/cli.cpp:1083–1086).
If that canary succeeds, crashes, times out, or lacks either marker, the build has not proved a clean fail-closed refusal and no positive result from it is admissible.
3. Pinning and building the runtime
git clone https://github.com/ggml-org/whisper.cpp.git
cd whisper.cpp
git checkout 97c56f1dc1d1100a9d859c865a20c82d22f823ed
git apply /path/to/the/exact-compatibility.patch
git diff --no-ext-diff --binary -- src/whisper.cpp
# output must byte-match the declared patch; status must be exactly:
# M src/whisper.cpp
cmake -B build-coreml \
-DWHISPER_COREML=1 \
-DWHISPER_COREML_ALLOW_FALLBACK=OFF
cmake --build build-coreml -j --config Release
Both options default OFF (CMakeLists.txt:92–93). -DWHISPER_COREML=1
locates the CoreML and Foundation frameworks, adds -DWHISPER_USE_COREML,
and compiles the whisper.coreml target with -fobjc-arc
(src/CMakeLists.txt:30–83). A missing framework is a hard
FATAL_ERROR, not a downgrade (src/CMakeLists.txt:39).
Build-hygiene rules that have actually bitten this class of gate:
- Use a separate build directory from the non-Core-ML build. A stale
build/containing a non-Core-MLwhisper-cliwill pass every transcript test and fail every load assertion — or worse, be silently reused. - Record
sha256of the producedwhisper-cli, not just the source SHA. - Re-verify immediately before and after the build that the checkout contains exactly the declared one-line patch and no other tracked or untracked change. A generically "clean" checkout is unpatched and unsupported; a generically dirty checkout is unverifiable.
4. The artifact contract
4.1. Filename derivation is mechanical and unforgiving
whisper_get_coreml_path_encoder() (src/whisper.cpp:3328–3345) transforms
the exact -m argument string:
- truncate at the last
.in the whole path; - if the trailing
-segment is exactly five characters and matches-q?_?, drop it; - append
-encoder.mlmodelc.
-m argument |
Derived Core ML path | Note |
|---|---|---|
models/ggml-base.en.bin |
models/ggml-base.en-encoder.mlmodelc |
Upstream convention. |
.../ggml-model.bin |
.../ggml-model-encoder.mlmodelc |
the reference implementation's exported name. |
.../ggml-base.en-q5_0.bin |
.../ggml-base.en-encoder.mlmodelc |
Quantized models share the FP encoder. |
.../ggml-model-q8_0.bin |
.../ggml-model-encoder.mlmodelc |
Same rule. |
.../ggml-tune-q1_5.bin |
.../ggml-tune-encoder.mlmodelc |
Rule is syntactic, not a quant-mode allowlist. Any -q?_? suffix is eaten. |
/data/v1.4/ggml-model (no extension) |
/data/v1-encoder.mlmodelc |
Step 1 finds the dot in the directory. Always give the model a .bin extension. |
Because derivation uses the argument string, the .mlmodelc must live beside
the model as the caller spelled it. Copying the .bin to a temp directory
without the .mlmodelc is the negative control of section 2.4 — and the
accidental version of it is a common false failure.
4.2. Tensor contract
| Boundary | Value | Source |
|---|---|---|
| Core ML input feature name | logmel_data |
src/coreml/whisper-encoder-impl.m:25, :29 |
| Core ML output feature name | output |
src/coreml/whisper-encoder-impl.m:180 |
Input MLMultiArray shape |
[1, n_mel, n_ctx] where n_ctx is the ggml mel width |
src/coreml/whisper-encoder.mm:55–62 |
| ggml mel tensor | ggml_new_tensor_2d(F32, 2*n_audio_ctx, n_mels) → 3000 x 80 |
src/whisper.cpp:1997 |
| Effective input | [1, 80, 3000] ([1, 128, 3000] for large-v3 class) |
derived from the two rows above |
| Input dtype | MLMultiArrayDataTypeFloat32 |
src/coreml/whisper-encoder.mm:58 |
| Expected output element count | n_audio_ctx * n_audio_state (e.g. 1500 * 384 = 576,000 for tiny) |
src/whisper.cpp:2022 |
| Output copy | memcpy(out, ...dataPointer, count * sizeof(float)) |
src/coreml/whisper-encoder.mm:67 |
| Compute units requested | MLComputeUnitsAll (hardcoded) |
src/coreml/whisper-encoder.mm:29 |
Three consequences follow directly from that memcpy and are worth stating
plainly, because none of them produce an error message:
- Float32 output is mandatory. The copy multiplies the element count by
sizeof(float). A Core ML model whose interface dtype is Float16 would be read past its buffer. (Upstream's--quantizesetscompute_precision=FLOAT16but leaves the IO type unspecified, so the interface stays fp32; a hand-rolled fp16-IO export would not.) - An oversized output overruns
embd_enc. There is no shape check on the runtime side. Validate input and output shapes on the Python side before the artifact is ever placed beside the model. - A failed prediction is silent. Both
initWithContentsOfURL:andpredictionFromLogmel_data:are called witherror:nil(src/coreml/whisper-encoder.mm:31,:65). A prediction failure yields a nil output,countof0, a zero-length copy, and an encoder state that was never written — producing empty or hallucinated text with a clean log and aCore ML model loadedline. This is the failure mode the load marker cannot catch. Only transcript parity catches it.
4.3. Compilation to .mlmodelc requires full Xcode
xcrun coremlc compile path/to/ggml-model-encoder.mlpackage path/to/output-dir/
# produces path/to/output-dir/ggml-model-encoder.mlmodelc
Locally verified on macOS 15.7.7: xcrun --find coremlc resolves to
/Applications/Xcode.app/Contents/Developer/Toolchains/XcodeDefault.xctoolchain/usr/bin/coremlc.
/Library/Developer/CommandLineTools was installed on the same host and
contains no coremlc. Command Line Tools alone are not sufficient; a CI
runner with only CLT will fail here, not at conversion.
coremlc names the output after the input basename, which is why the
.mlpackage must already be named ggml-<basename>-encoder.mlpackage. This is
the same mechanism upstream relies on before renaming
(models/generate-coreml-model.sh:31–33).
5. Producing the encoder: upstream script versus repo-owned converter
5.1. Current state of generate-coreml-model.sh
models/generate-coreml-model.sh (38 lines) calls
models/convert-whisper-to-coreml.py --encoder-only True --optimize-ane True
for named OpenAI models, or convert-h5-to-coreml.py under -h5 for a Hugging
Face directory, then compiles and renames.
Facts about the current converter, from source:
| Property | Current value | Source |
|---|---|---|
| Input shape | (1, hparams.n_mels, 3000) — fixed, no RangeDim, no EnumeratedShapes |
models/convert-whisper-to-coreml.py:260 |
| Backend | convert_to="mlprogram" |
:266 |
| Precision | FLOAT32 unless --quantize (which the shell script never passes) |
:270 |
| Compute units at convert time | ct.ComputeUnit.ALL |
:269 |
| SDPA | globally disabled via whisper.model.MultiHeadAttention.use_sdpa = False |
:31 |
| ANE reshaping | nn.Linear → 1x1 nn.Conv2d, LayerNormANE, 4D (B, C, 1, S) |
:34–:166 |
| Model allowlist | fixed list of OpenAI names; arbitrary names rejected | :307 |
5.2. Corrections to the historical guide
Three claims in coreml-hybrid-whispercpp.md do not match current source:
| Historical claim | Current reality |
|---|---|
"Uses ct.RangeDim(1, 3000) for variable audio lengths" |
The encoder input is a fixed (1, n_mels, 3000) tensor. There is no flexible shape anywhere in the current converter. |
"make WHISPER_COREML=1" |
The Makefile path is gone; CMake with -DWHISPER_COREML=1 is the documented build. |
"np.allclose(torch_out, coreml_out, atol=1e-2) # FP16 tolerance" |
Far too loose for a release gate, and max-absolute-error is the wrong statistic entirely. See section 7. |
5.3. The dependency trap in the upstream path
ane_transformers is required by convert-whisper-to-coreml.py:10. Its
install_requires is:
"torch>=1.10.0,<=1.11.0",
"coremltools>=5.2.0",
"transformers>=4.18.0",
"protobuf>=3.1.0,<=3.20.1",
It is distributed as an sdist only and its last release, 0.1.3, was
uploaded 2022-08-09 — no release in roughly four years. Installing it into this
repository's environment (torch==2.11.0, coremltools==9.0) is a resolver
conflict or a catastrophic downgrade. If the upstream reference conversion is
ever needed for comparison, build it in a throwaway isolated environment;
never in the project environment.
Independently, coremltools==9.0 declares _TORCH_MAX_VERSION = "2.7.0"
(coremltools/_deps/__init__.py:158) and only warns when exceeded:
Torch version 2.11.0 has not been tested with coremltools. You may run into
unexpected errors. Torch 2.7.0 is the most recent version that has been tested.
_warn_if_above_max_supported_version uses logger.warning, not an exception
(coremltools/_deps/__init__.py:33–39). Treat that warning as a recorded
build fact, not noise — it is the single most likely explanation offered for
unexplained numeric drift, and the implementationsitory has already tested it (Torch 2.7
did not change the FP32 outcome; see the debug notes).
The related _SKLEARN_MAX_VERSION = "1.5.1"
(coremltools/_deps/__init__.py:57) is the reason the reference implementation caps
scikit-learn: above it, coremltools disables its sklearn converter at import.
5.4. Why the reference implementation converts from Hugging Face directly
The upstream script loads openai-whisper checkpoints by name. A fine-tuned
Hugging Face artifact is not on that allowlist, and round-tripping through
convert-h5-to-coreml.py adds an unpinned second converter. the implementationsitory's
whisper_tuner/export/coreml.py instead traces the HF encoder directly and
must therefore reproduce the runtime contract itself:
- input named
logmel_data, fixed(1, n_mels, 3000), fp32; - output named
output,(1, 1500, hidden_size), fp32; mlprogram;- compiled to
ggml-model-encoder.mlmodelcbesideggml-model.bin.
The 3000-versus-1500 confusion is worth naming because it is the natural error
here: config.max_source_positions is 1500 (the encoder's output
length), while the required mel width is 3000. Hugging Face enforces this
and produces an exact error:
expected_seq_length = self.config.max_source_positions * self.conv1.stride[0] * self.conv2.stride[0]
if input_features.shape[-1] != expected_seq_length:
raise ValueError(
f"Whisper expects the mel input features to be of length {expected_seq_length}, "
f"but found {input_features.shape[-1]}. Make sure to pad the input mel features to {expected_seq_length}."
)
(transformers/models/whisper/modeling_whisper.py:612–616.) 1500 * 1 * 2 = 3000. Derive the contract from max_source_positions and the two conv
strides; never hardcode either number.
6. Driving the pinned runtime headlessly from Python
6.1. Stream and flag contract
- All library logs go to stderr. The default log callback is
fputs(text, stderr); fflush(stderr);(src/whisper.cpp:9186–9187).whisper-cli's own diagnostics also usefprintf(stderr, ...). Capture both streams and search the concatenation; do not search stdout alone. - Never pass
-np/--no-prints. It installscb_log_disable, an empty log callback (examples/cli/cli.cpp:1047–1048), which erases every Core ML line. A gate run with-npcannot pass and cannot fail honestly — it is simply blind. - Never pass
-ac/--audio-ctx(examples/cli/cli.cpp:168). It changes the ggml mel width to2 * audio_ctx, which no longer matches the encoder's fixed 3000-frame input. No guard exists in source; the result is the silent nil-prediction path of section 4.2. - Keep
-pat its default of 1.whisper_full_parallelcallswhisper_init_state(ctx)once per extra processor (src/whisper.cpp:7831–7833), so-p Nperforms N Core ML loads, emits N load markers, and holds N model instances. Fine for throughput, hostile to a deterministic gate. - Pin decoding.
--language <lang> --no-timestamps --temperature 0 --temperature-inc 0removes temperature-fallback nondeterminism from the transcript comparison. - Sanitize the child environment. Third-party subprocesses should not inherit Hugging Face or cloud credentials.
6.2. Exit codes
Verified from examples/cli/cli.cpp:
| Code | Meaning | Line |
|---|---|---|
0 |
Success | 1355 |
1 |
Argument parse failure, or missing @response file |
994, 1013 |
2 |
No input file specified | 1032 |
3 |
Unknown DTW preset, or whisper_init_from_file_with_params returned null |
1077, 1085 |
4 |
Grammar parse failure | 1106 |
10 |
whisper_full_parallel failed to process audio |
1320 |
3 is the Core-ML-relevant one. Do not treat "nonzero" as sufficient for the
negative control: require 3 and the failure marker, otherwise a typo in
the model path produces the same "expected failure".
Why unpatched exit 3 is intermittent — and never acceptable
The source intends to return 3, but the pinned upstream tree leaves
whisper_state.batch uninitialized:
whisper_state * state = new whisper_state;
// ... backend, KV-cache, alignment-mask, and Core ML setup ...
whisper_batch batch; // member declaration, no initializer
// Core ML required-load failure calls whisper_free_state(state) here
state->batch = whisper_batch_init(...); // not reached
whisper_free_state then unconditionally calls
whisper_batch_free(state->batch). That function conditionally frees five
pointer members, but the conditions themselves read indeterminate pointers.
The resulting invalid free explains the intermittent SIGABRT; the earlier
Metal-race explanation was wrong.
The supported patch is deliberately one line:
- whisper_batch batch;
+ whisper_batch batch = {};
Aggregate zero-initialization makes each pointer null before any early return.
The normal path later overwrites the member with whisper_batch_init, so the
successful initialization order is unchanged.
Measured A/B on August 1, 2026:
| Source | Unpatched | Exact zero-init patch |
|---|---|---|
pinned 97c56f1… |
22/30 exit 3; 8/30 SIGABRT |
100/100 exit 3 |
| then-current upstream master | 23/30 exit 3; 7/30 SIGABRT |
100/100 exit 3 |
Issue #2750 independently records the same Core ML failure endpoint and invalid free, but it does not establish the reference implementation's root-cause analysis. Issue #3807 is unrelated and must not be cited for this defect.
The release gate accepts only the clean, fully observed path: exact exit 3,
the Core ML load-failure marker, and the final CLI initialization-failure
marker. A crash proves memory unsafety, not fallback discipline. It is a hard
failure even if a marker was printed. Retrying is forbidden because it would
turn undefined behavior into probabilistic success.
The exporter and receipt validator share the classifier and intrinsic evidence
predicate from whisper_tuner.export.runtime_contracts; invalid exits, missing
markers, malformed evidence hashes, and relabeled crashes all fail closed.
6.3. Timeouts
The first positive run against a freshly compiled .mlmodelc triggers on-device ANE
compilation. The upstream README states this directly: the first run is slow
because the ANE service compiles the model to a device-specific format, and
later runs are faster.
Practical consequence for an automated gate:
- Give the first invocation a generous timeout (minutes, not seconds). Community reports of 20–50 s first loads for large encoders are common and plausible; treat any single published number as an estimate, not a bound.
- A positive-path warm-up invocation is allowed before the measured positive run; record both wall-clock values rather than gating on either. Never warm or retry the negative control—the first canary result is the evidence.
whisper_print_timingsreportsencode time = ... ms / N runs(src/whisper.cpp:4288), which is the honest place to observe the warm encoder cost — but it measures the whisper.cpp side of the call, and includes Core ML dispatch.
6.4. Compatibility-patch ownership
Every supported build binds three identities:
- upstream whisper.cpp commit;
- SHA-256 of the exact compatibility-patch bytes;
- SHA-256 of the built
whisper-clibinary.
The cache name includes the patch digest, the embedded bytes are hashed before
cache access, and the checkout is validated before and after the build. The
validator accepts exactly the declared diff on src/whisper.cpp and no other
tracked or untracked state. Direct tests cover hash drift, clean/unpatched and
modified trees, extra files, materialization, atomic publication, and staging
cleanup after a failed patch application.
Keep the patch until a proposed pin bump is inspected and physically tested.
Delete it only when the new upstream source initializes whisper_batch safely
on every early-return path, the unpatched candidate passes the repeated
negative control without a crash, ordinary GGML and Core ML positive paths
still pass, and manifests/cache identity are migrated deliberately. Current
upstream 2ca53bb… still declares whisper_batch batch;, so the deletion
condition is not met.
Sanitizer builds are useful upstream-quality evidence, but do not overclaim them: AddressSanitizer can catch the resulting invalid free, while detection of the uninitialized read itself may require MemorySanitizer or platform-specific instrumentation. The physical exact-exit/marker gate remains mandatory.
7. Parity measurement
7.1. Bind parity to the compiled artifact, at the runtime's compute units
The runtime loads a .mlmodelc with MLComputeUnitsAll
(src/coreml/whisper-encoder.mm:29, :31). Parity measured against an
.mlpackage, or at some other compute-unit setting, describes a configuration
that is not deployed. coremltools can load the exact compiled directory:
import coremltools as ct
encoder = ct.models.CompiledMLModel(
"/abs/path/to/ggml-model-encoder.mlmodelc", # the published bytes
compute_units=ct.ComputeUnit.ALL, # matches whisper-encoder.mm:29
)
candidate = encoder.predict({"logmel_data": mel_fp32})["output"]
CompiledMLModel takes "the path to a compiled model directory, ending in
.mlmodelc" and accepts the same ComputeUnit enum as MLModel
(coremltools/models/_compiled_model.py). the implementationsitory currently predicts
through ct.models.MLModel(<package>) inside _convert_and_predict() in
whisper_tuner/export/coreml.py; switching the parity prediction to the
compiled path is an evidence-binding improvement (same bytes, same compute
units as the deployed run), not necessarily a numerical one.
The same applies to the export's compute_units option: exporting or measuring
under cpu-only produces a number that whisper.cpp will never reproduce,
because whisper.cpp cannot be told to use CPU only without patching
whisper-encoder.mm.
7.2. The reference side
Feed the same mel to both sides. Reference is the frozen Hugging Face
encoder, fp32, eval mode, on CPU, under torch.inference_mode(), with
attn_implementation="eager" so the traced graph and the reference agree.
Compare against last_hidden_state shaped (1, 1500, hidden_size).
7.3. SNR is the right statistic, and Apple agrees
the reference implementation's frozen FP32 policy v2 is SNR ≥ 100 dB AND cosine ≥ 0.9999, with max-absolute-error and mean-absolute-error recorded but non-gating. That choice is methodologically consistent with coremltools' own converter test harness, which validates numeric agreement with SNR/PSNR rather than an absolute-error ceiling:
def compute_snr_and_psnr(x, y):
assert len(x) == len(y)
eps = 1e-5
eps2 = 1e-10
noise = x - y
noise_var = np.sum(noise**2) / len(noise)
signal_energy = np.sum(y**2) / len(y)
max_signal_energy = np.amax(y**2)
snr = 10 * np.log10((signal_energy + eps) / (noise_var + eps2))
psnr = 10 * np.log10((max_signal_energy + eps) / (noise_var + eps2))
return snr, psnr
(coremltools/converters/mil/testing_utils.py:867–877.)
Two convention differences from the implementationsitory's
signal_to_noise_ratio_db(reference, candidate) are worth recording so the
numbers are comparable:
- coremltools treats the second argument as the signal; the implementationsitory treats the first (the PyTorch reference) as the signal. At ≥ 100 dB the two are numerically indistinguishable; at low SNR they are not.
- coremltools divides by element count and adds
eps; the implementationsitory sums and handles the degenerate cases by clamping. The ratio is identical for same-length tensors.
Do not restate the policy version history here — it lives in
whisper_tuner/resources/evaluation_suites.json, including the recorded
failure of policy v1. The load-bearing rule for a new run is: the gate is
whatever the frozen resource says, and it is chosen before the measurement.
7.4. End-to-end WER methodology
Numeric parity on one probe is necessary, not sufficient. The end-to-end comparison must:
- run the source Hugging Face model on the frozen fixtures to produce the reference transcript;
- run the same pinned whisper.cpp binary on the same
ggml-*.binwithout the.mlmodelcpresent, producing the GGML-encoder transcript; - run it again with the
.mlmodelcpresent, producing the hybrid transcript, and assert the four load conditions of section 2.3; - compare normalized transcripts and report WER deltas for hybrid-vs-source and hybrid-vs-GGML.
Step 2 is what separates "Core ML changed the answer" from "conversion was
already wrong". It requires a fallback-enabled or separately built runtime,
or simply the non-Core-ML build of the same commit — do not enable
WHISPER_COREML_ALLOW_FALLBACK in the build used for step 3.
Persist hashes and bounded edit counts, not raw audio or raw transcripts, for any fixture that is not already public.
8. Observing actual compute placement
The load marker says nothing about ANE, GPU, or CPU. Three independent instruments, in decreasing order of evidentiary strength:
8.1. MLComputePlan (structural, per-operation)
from coremltools.models.compute_plan import MLComputePlan
from coremltools import ComputeUnit
plan = MLComputePlan.load_from_path(
"/abs/path/to/ggml-model-encoder.mlmodelc",
compute_units=ComputeUnit.ALL,
)
# MLModelStructureProgram is not subscriptable on coremltools 9.0; functions
# are reached through its `functions` mapping (verified on this repo's pin).
program = plan.model_structure.program
for op in program.functions["main"].block.operations:
usage = plan.get_compute_device_usage_for_mlprogram_operation(op)
cost = plan.get_estimated_cost_for_mlprogram_operation(op)
# usage.preferred_compute_device, usage.supported_compute_devices, cost.weight
load_from_pathtakes an.mlmodelcdirectory — the same artifact whisper.cpp loads (coremltools/models/compute_plan.py:403–448).MLComputePlanDeviceUsageexposespreferred_compute_deviceandsupported_compute_devices(compute_plan.py:295–310);MLComputePlanCost.weightis a[0.0, 1.0]share of total model execution (compute_plan.py:313–323).- Device classes are
MLCPUComputeDevice,MLGPUComputeDevice, andMLNeuralEngineComputeDevice(which also exposestotal_core_count) (coremltools/models/compute_device.py:68–135). - Availability: macOS 14.4+ / iOS 17.4+ (Apple's
MLComputePlanreference).load_from_pathraisesValueError("MLComputePlan is not supported.")when the native proxy is unavailable.
What it proves: the framework's plan. It is a strong, cheap, scriptable signal and it is per-operation, so it also identifies which ops break the ANE chain. It is not a recording of an actual run.
8.2. Instruments Core ML template (behavioral)
The Core ML template is present in a stock Xcode install — locally verified
via xcrun xctrace list templates on macOS 15.7.7. It can be driven headlessly:
xcrun xctrace record \
--template 'Core ML' \
--output coreml-run.trace \
--launch -- \
./build-coreml/bin/whisper-cli -m .../ggml-model.bin -f .../fixture.wav
What it proves: what actually executed, including per-model load and prediction
events. Cost: the .trace needs inspection, and automated extraction is more
work than MLComputePlan.
8.3. powermetrics --samplers ane_power (corroborating only)
Locally verified on macOS 15.7.7: powermetrics documents an ane_power
sampler — "dedicated rail ane power and frequency info" — and includes it in
both the default and all sampler groups. Nonzero ANE rail power during the
encoder window corroborates ANE execution.
Caveats that keep this out of the gate: it requires sudo, it is system-wide
rather than per-process, and Apple's own help text warns that "Average power
values reported by powermetrics are estimated and may be inaccurate".
8.4. The rule
Do not claim ANE residency from ComputeUnit.ALL, from a successful
transcription, from the load marker, or from a speedup. Claim it only from
MLComputePlan or an Instruments trace, and say which one.
9. Failure modes
| Symptom | Cause | Fix |
|---|---|---|
| No Core ML lines at all; run succeeds | Binary built without -DWHISPER_COREML=1, or a stale non-Core-ML whisper-cli was invoked |
Assert COREML = 1; build into a distinct directory; hash the binary |
failed to load Core ML model from '<path>', run continues |
Build enabled WHISPER_COREML_ALLOW_FALLBACK |
Rebuild with it OFF; assert the failure marker is absent |
failed to load ..., exit 3 |
.mlmodelc missing, misnamed, unreadable, or built for an incompatible OS |
Recompute the derivation of section 4.1 against the literal -m string |
| Load marker present, transcript empty or nonsense | Prediction failed silently (error:nil) — usually a feature-name or shape mismatch |
Validate input/output names and shapes on the Python side; compare against the GGML-encoder transcript |
Load marker present, garbage only with -ac N |
Mel width 2N ≠ the fixed 3000-frame encoder input |
Never pass -ac with a Core ML encoder |
Gate log has no Core ML lines but exit 0 |
-np / --no-prints silenced the log callback |
Remove -np |
| N load markers for one file | -p N created N states, each loading Core ML |
Use -p 1 for gates |
| First positive run appears hung | ANE compiling the model to a device-specific format | Raise the first-run timeout; add a positive-path warm-up run |
xcrun coremlc not found in CI |
Only Command Line Tools installed | Install full Xcode; xcrun --find coremlc must resolve inside Xcode.app |
ane_transformers install destroys the environment |
Its install_requires pins torch<=1.11.0 |
Never install it into the project env; isolate it, or use the implementation's HF converter |
| Derived path truncated at a directory name | Model file has no extension and a parent directory contains a . |
Always name the model *.bin |
10. What not to do
| Anti-pattern | Why it fails | Do this instead |
|---|---|---|
| Grep stdout for the load marker | All logs go to stderr | Concatenate both streams, or capture stderr specifically |
Use -np to keep gate output tidy |
Silences the exact lines being asserted | Keep the log; hash it into evidence |
| Accept "nonzero exit" as the negative control | Path typos and bad args also exit nonzero | Require exit 3 plus both failure markers |
Accept or retry SIGABRT |
Converts undefined behavior into probabilistic success | Fail immediately; patch the uninitialized batch and rerun once from a known identity |
| Require a clean whisper.cpp checkout | The supported runtime deliberately carries one patch | Require the exact declared diff and no other change |
Assert only COREML = 1 |
Proves the build, not the load | Assert both markers plus the absent-failure condition |
Assert only Core ML model loaded |
A fallback-enabled build could still have failed earlier in another run's log | Assert the full four-condition set from one run |
Measure parity on the .mlpackage |
Not the bytes the runtime loads | Use CompiledMLModel on the published .mlmodelc |
Measure parity under cpu-only |
whisper.cpp hardcodes MLComputeUnitsAll |
Match the runtime's compute units |
Gate on np.allclose(..., atol=1e-2) |
Tolerance chosen by folklore, statistic chosen wrong | Use the frozen SNR + cosine policy |
| Claim ANE from a 3x speedup | Metal GPU also produces large speedups | MLComputePlan or Instruments, named explicitly |
| Rebuild the artifact after the gate passes | The published bytes were never tested | Gate the exact digest that ships |
| Reuse one build directory for both runtimes | Silent ordinary-encoder runs | Separate build/ and build-coreml/ |
11. Verified versus plausible
Verified (primary source, cited by file and line above): every log string
and its emission condition; the ctx_coreml → whisper_encode_external proof
chain; WHISPER_COREML / WHISPER_COREML_ALLOW_FALLBACK defaults and effects;
the -encoder.mlmodelc derivation algorithm; stderr as the log sink; -np
silencing; whisper-cli exit codes; the [1, n_mels, 3000] input and
n_ctx * n_state float output with its unchecked memcpy; hardcoded
MLComputeUnitsAll; error:nil on both Core ML calls; per-processor state
creation under -p; upstream converter's fixed 3000 shape and FP32 default;
ane_transformers pins and its 2022 last release; coremltools 9.0
_TORCH_MAX_VERSION = 2.7.0 (warning only) and _SKLEARN_MAX_VERSION = 1.5.1;
compute_snr_and_psnr; MLComputePlan API surface and its macOS 14.4+
availability; CompiledMLModel accepting .mlmodelc; the Hugging Face
3000-frame check and its exact error text.
Verified locally on macOS 15.7.7 (2026-07-30): coremlc lives only inside
Xcode.app, not in Command Line Tools; Instruments ships a Core ML template;
powermetrics documents an ane_power sampler in its default and all
groups.
Verified locally on 2026-08-01: pinned and current upstream source leave
whisper_batch uninitialized across six pre-batch-init failure branches;
teardown unconditionally frees the member. The repeated unpatched/patch A/B in
section 6.2 produced 15 crashes in 60 unpatched attempts and zero crashes in
200 patched attempts. Dedicated ordinary-GGML and Core ML hybrid gates passed
with upstream, patch, and binary identities recorded. The exact physical nodes
must replay after the final checkout is frozen.
Plausible (community reports, not primary): first-load ANE compile times in
the tens of seconds for larger encoders; ANECompilerService pegging a core or
appearing to hang during that compile; the kill -9 ANECompilerService
workaround in the historical guide. These are consistent with upstream's own
"first run is slow" statement, but no primary source bounds the latency. Treat
any specific number as an observation from one machine.
Not established, and deliberately not claimed here: that the implementationsitory's recorded FP32 residual is caused by the torch-version warning, by compute-unit placement, or by any specific coremltools defect. The debug notes record that Torch 2.7, every compute-unit mode, and multiple deployment targets all failed the same way.
Works Cited
- whisper.cpp
src/whisper.cpp@97c56f1— Core ML load block, log strings, path derivation, external-encoder predicate, mel tensor, timings, default log sink. Retrieved 2026-07-30. - whisper.cpp
src/whisper.cpp@ current checkedmaster(2ca53bb) — the same uninitialized-batch lifecycle remains as of 2026-08-01. - whisper.cpp
src/coreml/whisper-encoder.mm—MLComputeUnitsAll,error:nil,MLMultiArrayshape/strides, outputmemcpy. Retrieved 2026-07-30. - whisper.cpp
src/coreml/whisper-encoder-impl.m/.h—logmel_data/outputfeature names, documented1 × 80 × 3000input. Retrieved 2026-07-30. - whisper.cpp
src/CMakeLists.txt—WHISPER_USE_COREMLdefinition, frameworkFATAL_ERROR,whisper.coremltarget. Retrieved 2026-07-30. - whisper.cpp
CMakeLists.txt—WHISPER_COREMLandWHISPER_COREML_ALLOW_FALLBACKdefault toOFF. Retrieved 2026-07-30. - whisper.cpp
examples/cli/cli.cpp— exit codes,system_infoprint,--no-printslog disable,--audio-ctx. Retrieved 2026-07-30. - whisper.cpp
models/generate-coreml-model.sh— conversion,xcrun coremlc compile, rename toggml-<name>-encoder.mlmodelc. Retrieved 2026-07-30. - whisper.cpp
models/convert-whisper-to-coreml.py— fixed(1, n_mels, 3000)input,mlprogram, FP32 default, ANE reshaping, SDPA disable. Retrieved 2026-07-30. - whisper.cpp README, "Core ML support" — documented build, expected log lines, first-run compile statement, Xcode and Python 3.11 guidance. Retrieved 2026-07-30.
- whisper.cpp issue #2278, "When built with CoreML, can no longer run normal models" — community-visible consequence of fail-closed init; open since 2024-07-03. Retrieved 2026-07-30.
- coremltools
coremltools/_deps/__init__.py@ 9.0 —_TORCH_MAX_VERSION,_SKLEARN_MAX_VERSION, warning-only enforcement. Retrieved 2026-07-30. - coremltools
coremltools/converters/mil/testing_utils.py@ 9.0 —compute_snr_and_psnr. Retrieved 2026-07-30. - coremltools
coremltools/models/compute_plan.py@ 9.0 —MLComputePlan.load_from_path,MLComputePlanDeviceUsage,MLComputePlanCost. Retrieved 2026-07-30. - coremltools
coremltools/models/compute_device.py@ 9.0 — CPU/GPU/Neural Engine device classes. Retrieved 2026-07-30. - coremltools
coremltools/models/_compiled_model.py@ 9.0 —CompiledMLModelloads an.mlmodelcwith a chosenComputeUnit. Retrieved 2026-07-30. - coremltools 9.0 release notes — published 2025-11-10; "Support for PyTorch 2.7"; Python 3.13 support;
macOS26targets. Retrieved 2026-07-30. - Apple Developer,
MLComputePlan— availability macOS 14.4+, iOS 17.4+. Retrieved 2026-07-30. - Hugging Face Transformers,
modeling_whisper.py—expected_seq_length = max_source_positions * conv1.stride * conv2.strideand its exactValueError. Retrieved 2026-07-30. apple/ml-ane-transformerssetup.py—torch>=1.10.0,<=1.11.0,protobuf<=3.20.1. Retrieved 2026-07-30.ane_transformerson PyPI — 0.1.3 uploaded 2022-08-09; sdist only. Retrieved 2026-07-30.- whisper.cpp issue #2750 — independent Core ML load-failure and invalid-free symptom report; it does not identify the root cause. Retrieved 2026-08-01.