The Collective Library
Full-length Apple Silicon reverse-engineering reports produced while solving real model-runtime, compiler, memory, and data-integrity failures. These public editions remove private paths, internal run identifiers, and campaign-control notes while preserving the technical work.
Batch-1 Transformer Decode on Apple Silicon
Transformer decode / latency. A per-token latency model for batch-1 autoregressive decode across CPU, Metal, Core ML state, and the Neural Engine.
Read the full report · Markdown source
Low-Memory Whisper Turbo Inference on Apple Silicon
Whisper / memory. Darwin memory accounting, duplicate encoder ownership, Metal buffers, Core ML residency, and the measurements that distinguish file size from peak footprint.
Read the full report · Markdown source
Physical whisper.cpp + Core ML Encoder Verification on Apple Silicon (2026)
whisper.cpp / physical proof. A physical-device contract for proving which encoder executed, binding parity to the exact artifact, and rejecting silent ordinary-encoder fallback.
Read the full report · Markdown source
iOS Device Benchmarking with devicectl
iPhone / devicectl. Build, install, stage, launch, inspect compute plans, capture crash evidence, and recover durable benchmark receipts from a physical iPhone.
Read the full report · Markdown source
A Developer's Reference Guide for Migrating to MLX and MLX Swift
PyTorch / MLX. Unified-memory semantics, lazy execution, device streams, parameter translation, numerical parity, and Swift deployment boundaries.
Read the full report · Markdown source
Fast Accuracy-Preserving RNNT Beam Search on Apple Silicon
RNNT / beam state machine. Hoisted projections, fixed-batch hypothesis graphs, exact blank and recombination semantics, and the latency experiment that preserves WER.
Read the full report · Markdown source
External-Encoder, Decoder-Only Whisper Runtime on Apple Silicon
Whisper / runtime ownership. How to remove duplicate native encoder weights when Core ML owns encoder compute without breaking model loading, state, or fallback behavior.
Read the full report · Markdown source
Core ML Preallocated Multi-Output Buffers on Apple Silicon
Core ML / allocation. What public Core ML APIs do and do not promise about caller-owned output backings, multi-array storage, copies, and measurable allocation wins.
Read the full report · Markdown source
Core ML Compute Unit Scheduling on Apple Silicon: An Advanced Engineering Field Guide
Core ML / heterogeneous scheduling. CPU, GPU, and ANE partitioning; silent fallback; compute-plan inspection; operator constraints; and device-side verification.
Read the full report · Markdown source
Detecting Metal GPU Faults in PyTorch, When PyTorch Tells You Nothing
Metal / data integrity. Why an aborted command buffer can yield plausible WER, which sentinel patterns survive tokenization, and where the fail-closed gate belongs.
Read the full report · Markdown source
Kokoro A14 iPhone Generator Execution Guide
A14 / ANE compiler. Tensor-axis ceilings, execution-plan failure, generator re-chunking, and the minimum fixed-shape probe that settles ANE feasibility.
Read the full report · Markdown source
Core ML Fused SDPA for Whisper-Shaped Encoders on Base M1
Attention / Core ML MIL. A fused-SDPA experiment that separates MIL spelling, output correctness, graph placement, and actual Neural Engine execution.
Read the full report · Markdown source
ANE Compile Cost and Compiled-Bundle Caching for Large Core ML Models
ANE / cold start. Why first load can take minutes, where compiled artifacts live, which cache layers matter, and how to distinguish compilation from inference.
Read the full report · Markdown source
Stateful Core ML Whisper Decoders on Apple Silicon
Whisper / stateful decode. Internal KV-cache state, update semantics, branching limits, placement falsifiers, and the evidence required before claiming a speedup.
Read the full report · Markdown source
Core ML Whisper Encoder Compression and ANE Graph Surgery
Whisper / ANE graph surgery. Low-bit weight storage, graph rewrites, package size, placement, numerical parity, and the physical benchmark needed to prove a useful compression.