Low-Memory Whisper Turbo Inference on Apple Silicon
Collective Library edition. This is the complete technical report. Private filesystem paths, internal run identifiers, campaign-control notes, and repository navigation were removed. Technical claims, code, measurements, evidence labels, citations, corrections, and falsification criteria are preserved.
August 12, 2026
Executive Summary
The minimum-memory architecture is not “mmap the model and tune Darwin page
hints.” The pinned whisper.cpp loader does not mmap the model at all. It opens
the custom GGML .bin with std::ifstream, allocates backend buffers for every
declared encoder and decoder tensor, and streams the weight bytes into those
buffers. When Core ML is enabled, a separate MLModel performs encoder
inference, but the native encoder weights remain allocated and loaded.
That establishes a deletion-first ladder:
- instrument the current resident FP16 hybrid and seal its real process and system footprint;
- run the stock whole-model Q5_0 control to learn the safe quantization and bandwidth effect;
- add an explicit fail-closed external-encoder load mode that never allocates or loads native encoder tensors;
- quantize only the remaining decoder-side weights after the lossless encoder-deletion control proves exact output parity;
- right-size one serialized state for the product's bounded audio and search contract; and
- compress the Core ML encoder only after the host-side waste is gone.
TASK_VM_INFO.phys_footprint is an important app-process metric, but it is not
the sole truth and macOS has no documented universal “jetsam threshold” for a
desktop process. A product receipt must pair the worker's footprint with RSS,
dirty/clean/wired VM categories, relevant service processes, system
compression and swap deltas, latency, and transcript parity.
The raw report's mandatory caller-side @autoreleasepool and implied SIGBUS
cure are also unsupported for this pin. The pinned Metal backend already wraps
graph execution and tensor-transfer paths in autorelease pools, retains and
releases command buffers explicitly, and the Core ML wrapper uses ARC plus a
prediction-local pool. Add another pool only when a measured application-level
Objective-C loop accumulates autoreleased objects; do not treat it as a generic
Metal crash fix.
1. Corrected Runtime Facts
1.1. Claims from the raw report that require correction
| Raw-report claim | Corrected verdict |
|---|---|
| Whisper weights are loaded through GGUF/GGML mmap | False for the pinned runtime. Whisper uses its custom sequential GGML .bin, opens it with std::ifstream, allocates backend tensors, and reads bytes into them. |
Clean mapped pages make model weights nearly free in phys_footprint |
Not applicable to the pinned loader. The weight buffers are allocated runtime storage, not a lazy read-only file mapping. |
phys_footprint is the only metric that predicts termination |
Overstated. It is the task's charged footprint ledger and a primary metric for dirty memory impact. Clean working set, wired memory, opaque service allocations, compressor activity, swap, and system pressure still matter. |
| Core ML creates exactly two physical copies of the encoder | Not established. The source proves a full native encoder allocation plus a separate Core ML model instance. Core ML's physical representation and accounting require measurement. |
mlock converts clean file pages into charged footprint |
Imprecise and irrelevant here. mlock keeps mapped pages resident and is subject to wired and resource limits; the pinned model is not file-mapped. |
MADV_WILLNEED is the recommended warm-model solution |
Not applicable here. madvise only advises an existing mapped range. The pinned model weights live in backend allocations. |
MADV_FREE safely leaves model pages available for reuse |
Dangerous. MADV_FREE says the application no longer needs the contents and permits immediate reuse. It is not a model-weight cache hint. |
Every inference loop needs a caller @autoreleasepool to avoid SIGBUS |
Unsupported. The pinned Metal and Core ML paths already contain scoped pools. SIGBUS has many causes; no primary source ties it to a missing outer pool here. |
| One context can serve parallel states with zero duplication | Contradicted by the public header. The C interface says the same whisper_context must not be used concurrently. Every state also owns caches, schedulers, backends, and a Core ML model handle. |
Setting audio_ctx=600 automatically shrinks cross-K/V allocation |
False for this pin. State initialization allocates cross and pad caches from the model's full n_audio_ctx, before request parameters set exp_n_audio_ctx. |
| Q5_0 has negligible accuracy loss and is the Apple Pareto frontier | Unverified for this product. Upstream supports Q5_0; accuracy and latency remain model-, corpus-, and device-specific promotion gates. |
1.2. The model file is not GGUF
The pinned Whisper loader expects its own ordered GGML binary. It checks
GGML_FILE_MAGIC, then reads hyperparameters, mel filters, vocabulary, and
named tensors. Vendored GGUF support elsewhere in the implementationsitory does not make
this load path GGUF.
This distinction matters for memory engineering. Advice written for a GGUF mmap loader cannot be transferred to this Whisper runtime without inspecting the actual loader.
1.3. What Core ML replaces—and what it does not
Core ML replaces encoder compute after state initialization:
audio
|
v
mel preparation in whisper.cpp
|
v
Core ML encoder MLModel --------> encoder embeddings
|
v
whisper.cpp cross cache
|
v
whisper.cpp decoder
It does not currently replace native encoder storage. The model loader
unconditionally constructs encoder tensor metadata, allocates all corresponding
backend buffers, and loads every tensor before whisper_init_state opens the
Core ML sidecar.
The source therefore proves structural redundancy: the native encoder tensor payload exists even though the native encoder graph is skipped. It does not prove that Core ML adds a byte-for-byte second FP16 physical copy. Core ML may compile, compress, cache, share, or charge its representation differently, and some allocations may reside in framework or service processes.
2. Allocation Topology
2.1. Model-global context
whisper_context owns:
- model hyperparameters, vocabulary, and mel filters;
- tensor metadata;
- backend buffers containing all loaded model weights; and
- optionally a default
whisper_statewhen a non-_no_stateinitializer is used.
The pinned loader's decisive sequence is:
- create every encoder and decoder tensor description;
- allocate backend buffers for those descriptions;
- read every weight tensor directly into a host-capable buffer, or through a temporary buffer into another backend; and
- reject unknown, missing, or shape-mismatched tensors.
The file is closed after load. Deleting or advising file-cache pages cannot release the model's allocated backend storage.
2.2. Per-state allocations
Every whisper_state owns substantial mutable state:
- backend instances;
- self-attention, cross-attention, and flash-attention pad caches;
- convolution, encoder, cross-attention, and decoder scheduler storage;
- mel, logits, batch, candidate, and result vectors;
- optional DTW and VAD structures; and
- under
WHISPER_USE_COREML, a separately initialized Core ML model handle.
When an external encoder is live, the pin omits the native encoder scheduler, but it still allocates the convolution shim, cross-attention scheduler, decoder scheduler, all caches, and all model weights.
Use whisper_init_from_file_with_params_no_state followed by exactly one
whisper_init_state when custom state ownership is required. Calling the
ordinary initializer and then allocating an additional state silently creates
two states.
2.3. One serialized state is the memory floor
The public header's contract says the same context must not be used by multiple
threads concurrently. whisper_full is explicitly not thread-safe for the same
context. Separate states expose finer-grained APIs, but the header does not
promote concurrent same-context execution as supported.
For a dictation product that finalizes one active speaker at a time, the minimum-memory topology is therefore simple:
one model context
|
+-- one state
|
+-- one serialized inference queue
Do not create a state pool until product concurrency requires it. Each state adds caches and schedulers and initializes another Core ML model object. Framework-internal sharing may reduce the incremental cost, but that is a measurement question, not a zero-duplication guarantee.
2.4. Metal shared buffers do not imply a CPU-plus-GPU copy
The pinned Metal backend can allocate CPU storage and wrap it with
newBufferWithBytesNoCopy using shared storage. In that path the MTLBuffer
references the same allocation; it is not evidence of a second full GPU copy.
The backend keeps the underlying allocation alive with the Metal buffer and
releases them together.
Other paths can use private Metal storage and explicit blits. Record the actual backend buffer names and storage path from the pinned build. “Unified memory” does not mean every API object aliases the same bytes, and “Metal buffer” does not by itself mean duplicate DRAM.
3. Exact State Arithmetic for Turbo
3.1. Architecture constants
The official openai/whisper-large-v3-turbo configuration and pinned GGML
header use:
- vocabulary: 51,866;
- audio context: 1,500 positions;
- model width: 1,280;
- encoder layers: 32;
- text context: 448 tokens; and
- decoder layers: 4.
The model's FP16 file and loader-reported weight size are recorded in the local campaign note. File size is artifact evidence; it is not a peak-memory measurement.
3.2. K/V allocations use padded maxima
The pinned state initializer uses FP16 intermediate caches and rounds context lengths to multiples of 256. For Turbo:
| Cache | Source dimensions | Allocated bytes |
|---|---|---|
| Self K and V | 4 layers * 512 text slots * 1280 width * 2 arrays * 2 bytes |
10,485,760 |
| Cross K and V | 4 layers * 1536 audio slots * 1280 width * 2 arrays * 2 bytes |
31,457,280 |
| Encoder pad K and V | 1 layer * 1536 audio slots * 1280 width * 2 arrays * 2 bytes |
7,864,320 |
| Total | All three caches | 49,807,360 |
This corrects two common mistakes:
- use the padded 512 and 1,536 allocations, not logical 448 and 1,500; and
- include the separate pad cache.
Setting request audio_ctx=600 changes graph views after state creation. It
does not change these allocations. A bounded-state patch would need to pass the
product's maximum audio positions into state initialization and allocate
GGML_PAD(max_audio_ctx, 256) deliberately.
That patch has a small ceiling compared with native encoder-weight deletion. At 600 positions, the rounded audio cache is 768 positions, so shrinking cross and pad caches can save only the difference between the 1,536- and 768-slot allocations. Prove it after the gigabyte-scale waste is removed.
3.3. Reserved address space is not physical footprint
The pin reserves n_vocab * n_text_ctx floats for the state's logits vector:
92,943,872 bytes for Turbo. std::vector::reserve obtains capacity, but Darwin
does not necessarily charge every untouched demand-zero page to resident or
physical footprint immediately. Runtime resize and tensor copies determine
how much becomes touched.
Do not claim a 92.9 MB physical saving merely by removing the reserve. Sample before and after model load, state creation, first prefill, steady greedy decoding, and teardown. If the pages remain untouched, deleting the reserve may reduce virtual size without materially reducing peak inference RAM.
3.4. Compute buffers are empirical
The runtime logs scheduler sizes for convolution, encoder, cross-attention, and decoder paths. Those values depend on backend, flash attention, graph shape, and build. Capture the log lines in the receipt instead of deriving them from parameter count.
An external encoder should remove the native encoder scheduler. It does not remove:
- the mel/conv shim;
- encoder output storage;
- cross-attention K/V construction;
- decoder compute storage; or
- native encoder weights in the current loader.
4. Darwin Memory Accounting
4.1. No single column is enough
The installed footprint(1) documentation defines a process footprint around
dirty memory charged to that process and recommends reducing dirty memory
first. It separately reports clean, reclaimable, swapped, and wired memory.
Clean file-backed working sets can still cause latency thrash; wired memory
cannot be paged, compressed, or reclaimed.
Use a metric set rather than a winner:
| Metric | What it answers | Important blind spot |
|---|---|---|
TASK_VM_INFO.phys_footprint |
Current footprint charged to the worker task | Does not attribute opaque work in other processes or describe clean working-set latency |
ledger_phys_footprint_peak |
Peak charged footprint observed for the worker | Requires a sufficiently new TASK_VM_INFO revision; still worker-only |
resident_size / RSS |
Pages currently resident for the process | Includes categories with different reclaimability and can overstate durable pressure |
footprint dirty/clean/reclaimable/wired columns |
Category-level task or multi-process accounting | Cross-process inspection can require root; sampling has overhead |
vmmap -summary |
VM regions, dirty/swapped allocation, and mapped-file categories | Snapshot, not a peak; permissions may restrict inspection |
vm_stat deltas |
System compressor, page-in/out, and swap activity | System-wide and sensitive to unrelated work |
memory_pressure snapshot |
Current system memory statistics and free percentage | Not per-process; -S simulates pressure and invalidates a benchmark |
| Core ML/ANE service census | Whether work appears outside the worker | Process names and accounting can change by OS version; causality needs a control |
macOS desktop process termination is not specified as one public footprint threshold. A low worker footprint can coexist with high system pressure, and a high RSS can include reclaimable or clean pages. Report both.
4.2. In-process sampler
The installed SDK exposes current and peak fields through TASK_VM_INFO.
Sample in-process so the worker can bind memory points to exact inference
events:
#include <mach/mach.h>
#include <mach/task_info.h>
#include <cstdint>
struct DarwinTaskMemory {
uint64_t resident_bytes;
uint64_t resident_peak_bytes;
uint64_t compressed_bytes;
uint64_t footprint_bytes;
uint64_t footprint_peak_bytes;
};
static bool read_darwin_task_memory(DarwinTaskMemory & out) {
task_vm_info_data_t info = {};
mach_msg_type_number_t count = TASK_VM_INFO_COUNT;
const kern_return_t result = task_info(
mach_task_self(),
TASK_VM_INFO,
reinterpret_cast<task_info_t>(&info),
&count);
if (result != KERN_SUCCESS || count < TASK_VM_INFO_REV1_COUNT) {
return false;
}
out.resident_bytes = info.resident_size;
out.resident_peak_bytes = info.resident_size_peak;
out.compressed_bytes = info.compressed;
out.footprint_bytes = info.phys_footprint;
out.footprint_peak_bytes =
count >= TASK_VM_INFO_REV3_COUNT &&
info.ledger_phys_footprint_peak > 0
? static_cast<uint64_t>(info.ledger_phys_footprint_peak)
: 0;
return true;
}
Treat a zero peak as unavailable unless the surrounding API call failed. Do not silently replace it with the current value.
4.3. Required measurement phases
For every candidate arm, capture the same event sequence:
- quiet-host admission and system baseline;
- worker process started, before model load;
- after native model context load;
- after the single state and Core ML model load;
- after first specialization or warmup;
- immediately before each measured request;
- at a high enough cadence during the request to observe the peak;
- immediately after final text;
- after an idle interval; and
- after explicit state/context teardown.
A before/after snapshot is insufficient when an intermediate compute buffer is the peak. Prefer in-process phase markers plus a 50-100 ms external sampler. Record sampler overhead with a no-sampler latency control.
4.4. Process and system scope
At minimum, preserve:
# Per-process VM category snapshot.
vmmap -summary "$WORKER_PID"
# Current RSS in KiB; not a peak and not a footprint substitute.
ps -o pid=,rss=,vsz=,command= -p "$WORKER_PID"
# System totals followed by one-second deltas.
vm_stat 1
# Current system memory statistics. Do not add -S to a measurement run.
memory_pressure
Use footprint --sample when privileges and overhead are acceptable. When
comparing multiple related processes, use its multi-process de-duplication
rather than summing RSS columns blindly.
Core ML may perform work in framework or service processes that are not child processes of the app. Record a process census before and during the Core ML arm and pair it with a native-encoder control. Do not hard-code a service name as a stable API or attribute every system-service delta to the model.
5. The Deletion-First Optimization Ladder
5.1. Stage 0: seal the existing FP16 hybrid
Before changing allocation, measure the exact current product candidate. The receipt must bind:
- model file SHA-256 and byte size;
- Core ML package and compiled-tree identity;
- whisper.cpp commit and dirty-source state;
- build flags, linked libraries, and compute-unit policy;
- process and system memory series;
- runtime allocation log lines;
- p50, p95, and maximum stop-to-final latency; and
- exact tokens, EOT, final text, and quality-slice results.
This stage converts “approximately 2-3 GB” into a physical baseline. Without it, later deltas cannot be attributed.
5.2. Stage 1: stock whole-model Q5_0
The pinned quantizer supports Q5_0 and applies it to eligible two-dimensional tensors. It deliberately skips encoder convolution biases, encoder positional embeddings, and decoder positional embeddings; non-two-dimensional tensors are also retained.
Q5_0 is a high-leverage first control because the loader stores weights in their quantized tensor type. It can reduce the large native backend buffer and memory bandwidth without a custom runtime patch.
It is not lossless. Require:
- artifact/header validation and exact source binding;
- model-load success with the same Core ML sidecar;
- latency and memory measured with the same protocol;
- exact-output diagnostics on frozen rows; and
- full WER and product-pathology qualification before promotion.
Do not state an expected byte size or accuracy delta before the generated artifact and bakeoff exist.
5.3. Stage 2: skip native encoder allocation under fail-closed Core ML
This is the largest source-verified deletion opportunity. The lossless control should retain:
- the original header, vocabulary, filters, and decoder tensors;
- decoder embeddings and output projection;
- all decoder self- and cross-attention weights;
- the external encoder output ABI; and
- a hard requirement that the exact Core ML model loads successfully.
It should omit backend allocation for:
- encoder positional embeddings;
- both encoder convolution layers; and
- all 32 native encoder transformer layers.
Do not begin by inventing a new “decoder-only GGML” container. The smaller first patch can keep the existing full artifact and teach an explicit external-encoder-only loader mode to validate encoder tensor names and shapes while seeking or discarding their bytes without allocating backend tensors. That isolates RAM deletion from file-format surgery.
The current generic loader callback has read, eof, and close, but no
seek/skip callback. Reading a skipped giant tensor into a temporary vector
would recreate the peak being removed. Add either:
- a bounded chunk discard path; or
- an optional checked
skipcallback implemented by the file loader.
The skip path must reject truncated input, overflow, unexpected tensor names,
wrong shapes, wrong types, missing tensors, duplicate tensors, and a Core ML
load failure. It must never silently fall back to native encoding because the
native weights do not exist.
The external encoder call must fail closed too. A successful model load is not
enough: propagate predictionFromFeatures errors, require exactly one Float32
[1, 600, 1280] output, validate its element count and strides before copying,
and return an encoding failure rather than decoding stale or malformed state.
An unchecked memcpy(output.count) is incompatible with a runtime whose native
encoder fallback has deliberately been deleted.
Promotion gate: exact token IDs, EOT, text, logits within a preregistered tolerance if captured, and latency parity against the full FP16 hybrid. The memory delta is then causally attributable to unused native encoder storage.
5.4. Stage 3: quantize the retained decoder
Only after Stage 2 proves lossless deletion should quantization be applied to the remaining eligible decoder tensors. This cleanly separates:
- exact-output changes caused by deleting unused tensors—which should be none; from
- numerical changes caused by decoder quantization.
A compact decoder-only shipping artifact can follow later if disk size and load I/O justify a new format. The runtime patch and its validation rules are the hard part; truncating bytes first merely creates an incompatible file.
5.5. Stage 4: right-size mutable state
After weight waste is gone, test smaller allocations in descending value:
- exactly one state and one serialized request;
- greedy search with one decoder rather than beam candidates;
- DTW/token-timestamp structures disabled when the product does not consume them;
- bounded cross and pad caches created from the product's maximum external encoder positions; and
- logits capacity and decoder work buffers sized from observed maximum batch rather than theoretical text context.
Every bound needs a fail-closed input check. A 600-position state must reject an encoder output longer than 600 rather than overrun, silently truncate, or reuse stale cache tails.
5.6. Stage 5: compress Core ML last
Core ML palettization can reduce package storage and may reduce runtime weight cost, but package bytes are not resident bytes. The compiler can transform or expand representations, and compute placement can change.
For each compression arm, separately prove:
- portable package size and compiled-tree size;
- numerical parity against the dense encoder;
- end-to-end transcript and WER behavior;
- actual ANE placement;
- cold specialization and warm-load behavior;
- app and service-process memory; and
- stop-to-final latency under the same contention contract.
Do not combine decoder quantization, native-encoder deletion, and Core ML compression in the first experiment. One changed mechanism per arm preserves causal evidence.
5.7. Idle residency is a product policy, not a kernel trick
Loading while the user speaks can hide resident initialization from stop-to-final latency. Releasing the context after an idle timeout can lower idle RAM. Neither reduces peak inference memory.
Measure three policies:
| Policy | User benefit | Memory cost | Risk |
|---|---|---|---|
| Always resident | Lowest request-start variance | Highest idle RAM | System pressure while unused |
| Load on recording start | Hides load behind speech | Peak unchanged, idle lower | Short utterances may stop before load finishes |
| Idle timeout | Amortizes nearby requests | Tunable | First request after timeout pays load/specialization |
Core ML's first device specialization is a separate product event. Do not confuse install-time compilation, warm model load, and steady prediction.
6. mmap, mlock, and madvise
6.1. Why they are not current levers
The pinned Whisper path does not expose a mapped model-weight range. Therefore:
- there is no model mapping to lock;
- there is no model mapping to prefetch with
MADV_WILLNEED; - dropping the source file's cache does not free backend weights; and
- a benchmark of file-cache advice would test I/O, not resident inference RAM.
Delete this work from the immediate plan.
6.2. If a future loader introduces mmap
Treat these APIs according to their platform contracts:
mlockkeeps the addressed pages resident until unlocked and is bounded by process and system limits. It increases non-reclaimable pressure and should be an explicit latency-versus-memory experiment, not a default.MADV_WILLNEEDis advisory. The man page does not promise asynchronous completion, full prefetch, or future residency.MADV_DONTNEEDsays the range is not expected soon; it is not a durability or latency guarantee.MADV_FREEpermits the VM to reuse contents. Never apply it to live model weights.
An mmap rewrite might reduce dirty weight storage by relying on clean file-backed pages, but Metal compatibility, tensor alignment, quantized access, page-fault tails, and Core ML coexistence all need a separate design and physical falsifier. It is not prerequisite to removing unused encoder tensors.
7. Autorelease Pools, No-Copy Buffers, and SIGBUS
7.1. What the pin already does
The pinned sources establish:
src/coreml/whisper-encoder.mmrefuses to compile without ARC;- Core ML prediction is wrapped in
@autoreleasepool; - Metal graph compute and tensor transfer paths contain their own
@autoreleasepoolblocks; - command buffers created with unretained references are explicitly retained, replaced, synchronized, and released by the backend; and
- shared Metal buffers can wrap backend-owned CPU storage with a nil deallocator while that storage remains owned by the backend buffer.
That is not evidence of a pool leak.
7.2. What an outer pool can and cannot do
An application-level @autoreleasepool around a long Objective-C loop can
bound objects created by the application or frameworks and released by
autorelease. Use it when an allocation trace shows growth between pool drains.
It cannot repair:
- a CPU pointer freed before an asynchronous GPU operation completes;
- out-of-bounds tensor views or incorrect strides;
- a command buffer or resource that is explicitly retained and never released;
- invalid model bytes;
- an OS or driver fault; or
- memory pressure caused by live, required buffers.
7.3. No-copy lifetime rule
Apple's no-copy Metal initializer wraps caller-provided bytes. The caller must
keep those bytes valid for the lifetime and use of the MTLBuffer. In an
asynchronous path, lifetime must extend through GPU completion, not merely the
encoding call.
When SIGBUS occurs, preserve the crash report, fault address, GPU error state, command-buffer status, tensor name and range, buffer ownership, and completion ordering. Do not diagnose “missing autorelease pool” from the signal alone.
8. Decision Framework
| Mechanism | Expected memory target | Accuracy risk | Latency risk | Evidence status |
|---|---|---|---|---|
| Measure current FP16 hybrid | None; establishes truth | None | Sampling overhead | Required first |
| Whole-model Q5_0 | Native backend weights | Real; must qualify | May improve bandwidth or add dequant cost | Source-supported candidate |
| Skip native encoder allocation | Unused native encoder weights | None if external path is exact and fail-closed | Load path changes; inference should match | Highest-value hypothesis from source |
| Quantize retained decoder | Decoder weights and bandwidth | Real; must qualify | Device-dependent | Test after lossless deletion |
Single _no_state context plus one state |
Duplicate caches, schedulers, Core ML objects | None for serialized product | Removes concurrency | Source-verified topology |
| Bound cross/pad cache to 600 | Tens of MB, not GB | Input-bound failure if contract violated | Small possible gain | Secondary source-supported patch |
| Remove oversized logits reserve | Virtual capacity; physical saving uncertain | Possible batch regression | Probably small | Measure before implementation |
| Disable DTW and beam search | Optional mutable state | Product-feature and decode-quality tradeoff | Usually favorable | Product-contract dependent |
| LUT-compress Core ML encoder | Core ML package and possibly runtime weights | Numerical and placement risk | Compile and prediction risk | Last-stage physical experiment |
| Idle unload | Idle memory only | None | Reload/specialization tail | Product policy |
| Add model mmap | Dirty weight storage, possibly | Loader/backend rewrite risk | Page-fault tails | Separate future program |
Add mlock/madvise now |
None in current loader | None | Engineering distraction | Delete |
9. Failure Modes and Debugging Playbook
9.1. File size falls but peak footprint does not
Likely causes:
- the quantized file still expands into another backend type;
- Core ML or compute buffers dominate the peak;
- the sampler missed the true peak;
- both old and new contexts overlap during model replacement; or
- allocator high-water pages remain charged after a transient conversion.
Check backend buffer log sizes, phase markers, context overlap, and
footprint/vmmap categories before changing quantization again.
9.2. Worker footprint falls but system pressure does not
Run the native-encoder and Core ML controls with identical host admission. Inspect relevant service-process and system deltas. Core ML work may be charged outside the worker, but unrelated services can move at the same time. Require repeatable paired deltas before attribution.
9.3. A second state causes a large jump
Verify the initializer. If the ordinary context initializer created a default
state and the application then called whisper_init_state, free the duplicate
and switch to the _no_state initializer. Next compare state creation with and
without Core ML to separate caches/schedulers from framework model residency.
9.4. audio_ctx=600 did not shrink initialization logs
This is expected for the pin. The request parameter is assigned after state creation, while K/V caches were allocated from full model hyperparameters. Implement a checked state-init bound or stop claiming a memory saving.
9.5. RSS is much larger than footprint
Use footprint and vmmap to split dirty, clean, reclaimable, wired, and
shared categories. Do not optimize the numerical difference itself. Optimize
the category that limits the product: charged pressure, clean working-set
thrash, wired allocations, or latency after eviction.
9.6. Memory grows across repeated requests
Separate three patterns:
- one-time lazy allocation that plateaus;
- allocator caching that remains reusable; and
- unbounded growth proportional to request count.
Run enough iterations to distinguish them, add explicit pool-drain markers at the application boundary, and record live object/VM categories. A stable higher plateau is not automatically a leak; linear growth is not automatically an autorelease problem.
9.7. Q5 changes one proper noun or number
That is an accuracy failure for exact-output parity even if aggregate WER barely moves. Keep the full pathology matrix: names, numbers, negations, punctuation, false starts, repetitions, silence, and clean dictation. Promote using the product's quality policy, not file size or average WER alone.
9.8. Metal faults after a no-copy change
Audit ownership and ordering first:
- identify the exact pointer and byte range;
- prove alignment and buffer bounds;
- prove the CPU allocation remains alive;
- wait for or otherwise order GPU completion before reuse/free;
- inspect command-buffer error details; and
- rerun with a copied-buffer control.
An outer autorelease pool is not a substitute for those proofs.
10. Production Qualification Receipt
Every memory candidate should emit one sealed record with:
Artifact identity
- source model and revision;
- GGML model SHA-256, byte size, and ftype;
- Core ML package and compiled-tree digests;
- whisper.cpp commit, patch digest, and dirty status;
- executable and linked-runtime digests; and
- OS, hardware model, RAM capacity, and boot identity.
Runtime topology
- context initializer used;
- state count;
- Core ML model count;
- backend buffer names and logged sizes;
- compute-unit setting;
- flash-attention, timestamp, VAD, search, and audio-context settings; and
- serialized or concurrent request policy.
Memory series
- worker current and peak
phys_footprint; - worker current and peak resident size;
- worker compressed bytes;
footprintcategory summaries when available;- relevant service-process series or an explicit “not attributable” marker;
- system wired/compressor/page-in/page-out/swap deltas; and
- baseline, post-load, post-state, warmup, per-request peak, idle, and teardown phase timestamps.
Product outcomes
- stop-to-final p50, p95, and maximum;
- cold load and first-specialization time reported separately;
- exact token IDs, EOT, and final text for diagnostic fixtures;
- WER and pathology gates on frozen slices;
- silence false alarms;
- memory growth slope over a long repeated-request soak; and
- cancellation, fallback, and idle-unload behavior.
The promotion rule is conjunctive: lower memory is useful only if latency, quality, stability, and fail-closed behavior still pass.
11. Anti-Patterns
| Anti-pattern | Why it fails | Use instead |
|---|---|---|
| Quote model file size as RAM | Ignores allocator, runtime, state, services, and compression | Phase-aligned physical measurements |
| Treat RSS as the sole metric | Mixes categories with different reclaimability | Footprint plus VM categories and system deltas |
Treat phys_footprint as the sole metric |
Misses clean working-set latency and other processes | Worker, multi-process, and system view |
Run memory_pressure -S during a benchmark |
It simulates pressure rather than observing it | Plain memory_pressure snapshot and vm_stat deltas |
Tune madvise for the pinned model |
There is no mapped weight range | Delete unused allocations first |
mlock the whole model |
Increases non-reclaimable residency and can hit limits | Keep resident only what the product needs |
| Prune the GGML file before changing the loader | Current loader requires every expected tensor | Add explicit external-only load semantics first |
| Enable Core ML fallback after skipping encoder weights | Fallback has no weights to execute | Fail closed at initialization |
| Allocate a default state plus a state pool | Duplicates mutable runtime objects | _no_state initializer plus the minimum state count |
| Assume Q5 preserves quality | Quantization changes decoder arithmetic | Frozen exact-output and WER/pathology bakeoff |
| Add caller pools until SIGBUS disappears | Masks evidence and does not prove ownership | Crash, buffer-lifetime, and command-status diagnosis |
| Combine three compression mechanisms in one arm | Makes deltas and regressions unattributable | One changed mechanism per sealed arm |
Works Cited
- Pinned whisper.cpp model loader and state implementation — custom GGML loading, backend weight allocation, external encoder selection, state caches, schedulers, and teardown.
- Pinned whisper.cpp public C API — context/state initializers, ownership, and same-context thread-safety boundary.
- Pinned whisper.cpp Core ML wrapper — ARC requirement,
MLModellifetime, compute units, no-copyMLMultiArrayinput, and prediction pool. - Pinned GGML Metal context — graph and transfer pools plus command-buffer ownership.
- Pinned GGML Metal device and buffer implementation — shared/private storage, no-copy wrapping, blits, and buffer teardown.
- Pinned whisper.cpp quantizer — supported model layout and explicit skip list.
- Pinned common GGML quantization implementation — supported types and two-dimensional-tensor rule.
- Official Whisper large-v3-turbo configuration — architecture dimensions used in allocation arithmetic.
- Apple XNU
task_vm_info— resident, compressed, footprint, peak, and neural ledger fields. The exact verified declarations also ship in the installed macOS 26.5 SDK. - Apple Metal no-copy buffer API — caller-provided storage contract.
- Installed macOS 26.5
footprint(1),vmmap(1),memory_pressure(1),vm_stat(1),mlock(2), andmadvise(2)man pages — local primary platform contracts used to correct the raw report's accounting and VM advice.