# The Collective Library

Full-length Apple Silicon reverse-engineering reports produced while solving real model-runtime, compiler, memory, and data-integrity failures. These public editions remove private paths, internal run identifiers, and campaign-control notes while preserving the technical work.

## [Batch-1 Transformer Decode on Apple Silicon](/reports/batch-1-transformer-decode-apple-silicon)

**Transformer decode / latency.** A per-token latency model for batch-1 autoregressive decode across CPU, Metal, Core ML state, and the Neural Engine.

[Read the full report](/reports/batch-1-transformer-decode-apple-silicon) · [Markdown source](/reports/batch-1-transformer-decode-apple-silicon.md)

## [Low-Memory Whisper Turbo Inference on Apple Silicon](/reports/low-memory-whisper-apple-silicon)

**Whisper / memory.** Darwin memory accounting, duplicate encoder ownership, Metal buffers, Core ML residency, and the measurements that distinguish file size from peak footprint.

[Read the full report](/reports/low-memory-whisper-apple-silicon) · [Markdown source](/reports/low-memory-whisper-apple-silicon.md)

## [Physical whisper.cpp + Core ML Encoder Verification on Apple Silicon (2026)](/reports/whispercpp-coreml-physical-verification)

**whisper.cpp / physical proof.** A physical-device contract for proving which encoder executed, binding parity to the exact artifact, and rejecting silent ordinary-encoder fallback.

[Read the full report](/reports/whispercpp-coreml-physical-verification) · [Markdown source](/reports/whispercpp-coreml-physical-verification.md)

## [iOS Device Benchmarking with devicectl](/reports/ios-device-benchmarking-devicectl)

**iPhone / devicectl.** Build, install, stage, launch, inspect compute plans, capture crash evidence, and recover durable benchmark receipts from a physical iPhone.

[Read the full report](/reports/ios-device-benchmarking-devicectl) · [Markdown source](/reports/ios-device-benchmarking-devicectl.md)

## [A Developer's Reference Guide for Migrating to MLX and MLX Swift](/reports/pytorch-to-mlx-migration)

**PyTorch / MLX.** Unified-memory semantics, lazy execution, device streams, parameter translation, numerical parity, and Swift deployment boundaries.

[Read the full report](/reports/pytorch-to-mlx-migration) · [Markdown source](/reports/pytorch-to-mlx-migration.md)

## [Fast Accuracy-Preserving RNNT Beam Search on Apple Silicon](/reports/rnnt-beam-search-apple-silicon)

**RNNT / beam state machine.** Hoisted projections, fixed-batch hypothesis graphs, exact blank and recombination semantics, and the latency experiment that preserves WER.

[Read the full report](/reports/rnnt-beam-search-apple-silicon) · [Markdown source](/reports/rnnt-beam-search-apple-silicon.md)

## [External-Encoder, Decoder-Only Whisper Runtime on Apple Silicon](/reports/whisper-external-coreml-encoder-runtime)

**Whisper / runtime ownership.** How to remove duplicate native encoder weights when Core ML owns encoder compute without breaking model loading, state, or fallback behavior.

[Read the full report](/reports/whisper-external-coreml-encoder-runtime) · [Markdown source](/reports/whisper-external-coreml-encoder-runtime.md)

## [Core ML Preallocated Multi-Output Buffers on Apple Silicon](/reports/coreml-preallocated-multi-output-buffers)

**Core ML / allocation.** What public Core ML APIs do and do not promise about caller-owned output backings, multi-array storage, copies, and measurable allocation wins.

[Read the full report](/reports/coreml-preallocated-multi-output-buffers) · [Markdown source](/reports/coreml-preallocated-multi-output-buffers.md)

## [Core ML Compute Unit Scheduling on Apple Silicon: An Advanced Engineering Field Guide](/reports/coreml-compute-unit-scheduling)

**Core ML / heterogeneous scheduling.** CPU, GPU, and ANE partitioning; silent fallback; compute-plan inspection; operator constraints; and device-side verification.

[Read the full report](/reports/coreml-compute-unit-scheduling) · [Markdown source](/reports/coreml-compute-unit-scheduling.md)

## [Detecting Metal GPU Faults in PyTorch, When PyTorch Tells You Nothing](/reports/metal-gpu-fault-detection-pytorch)

**Metal / data integrity.** Why an aborted command buffer can yield plausible WER, which sentinel patterns survive tokenization, and where the fail-closed gate belongs.

[Read the full report](/reports/metal-gpu-fault-detection-pytorch) · [Markdown source](/reports/metal-gpu-fault-detection-pytorch.md)

## [Kokoro A14 iPhone Generator Execution Guide](/reports/a14-coreml-generator-execution)

**A14 / ANE compiler.** Tensor-axis ceilings, execution-plan failure, generator re-chunking, and the minimum fixed-shape probe that settles ANE feasibility.

[Read the full report](/reports/a14-coreml-generator-execution) · [Markdown source](/reports/a14-coreml-generator-execution.md)

## [Core ML Fused SDPA for Whisper-Shaped Encoders on Base M1](/reports/coreml-fused-sdpa-base-m1)

**Attention / Core ML MIL.** A fused-SDPA experiment that separates MIL spelling, output correctness, graph placement, and actual Neural Engine execution.

[Read the full report](/reports/coreml-fused-sdpa-base-m1) · [Markdown source](/reports/coreml-fused-sdpa-base-m1.md)

## [ANE Compile Cost and Compiled-Bundle Caching for Large Core ML Models](/reports/ane-compile-cost-bundle-cache)

**ANE / cold start.** Why first load can take minutes, where compiled artifacts live, which cache layers matter, and how to distinguish compilation from inference.

[Read the full report](/reports/ane-compile-cost-bundle-cache) · [Markdown source](/reports/ane-compile-cost-bundle-cache.md)

## [Stateful Core ML Whisper Decoders on Apple Silicon](/reports/coreml-stateful-whisper-decoder)

**Whisper / stateful decode.** Internal KV-cache state, update semantics, branching limits, placement falsifiers, and the evidence required before claiming a speedup.

[Read the full report](/reports/coreml-stateful-whisper-decoder) · [Markdown source](/reports/coreml-stateful-whisper-decoder.md)

## [Core ML Whisper Encoder Compression and ANE Graph Surgery](/reports/coreml-whisper-encoder-compression)

**Whisper / ANE graph surgery.** Low-bit weight storage, graph rewrites, package size, placement, numerical parity, and the physical benchmark needed to prove a useful compression.

[Read the full report](/reports/coreml-whisper-encoder-compression) · [Markdown source](/reports/coreml-whisper-encoder-compression.md)
