The neural-cost v0.2.0 GPU benchmark moves from CPU roofline analysis to Apple Silicon's MPS GPU backend — measuring PyTorch and JAX across five architectures, four batch sizes, and both eager and compiled execution modes.
Writing
Notes from the workbench
Mostly about low-level systems, machine learning that has to survive contact with reality, and the constraints that made a decision obvious in hindsight.
Subscribe via RSSAn empirical investigation comparing PyTorch 2.14, JAX 0.11, and TensorFlow across five model architectures and four batch sizes, measuring compiler speedups, kernel overheads, and hardware limits on Apple Silicon.
A framework-neutral Python library that estimates a neural network's FLOPs and tensor traffic, measures its runtime, and uses a roofline model to tell you exactly where performance is being left on the table.
The old site was a single 41,000-character index.html. Every change was archaeology. Here is what I replaced it with and why the constraint was maintenance, not design.