✦ For YouGeopoliticsTechFinanceHealthEnergySportsCulture◆ SN Last Week★ Saved

Apple Silicon VM Llama.cpp Hits 11–16× via Metal Shim

By Hacker News · Summarized & edited by · 2026-08-11
Apple Silicon VM Llama.cpp Hits 11–16× via Metal Shim

Get the Tech newsletter

Daily tech — startups, AI labs, chips, the launches that shape the next decade. Free.

Why it matters: Apple Silicon developers running local LLMs inside macOS VMs have been capped well below host speed because Apple's paravirtualized driver reports an older GPU profile by default — Cua's shim lifts that ceiling, hitting 98% of bare-metal prompt throughput on TinyLlama and 94.82% on Gemma 4 12B on the tested M1 Ultra. The gains explicitly vary by host GPU, guest version, and workload, and Apple retains the right to alter the private guest Metal behavior the technique exploits in future macOS releases.

Share this story

Ask SkimNews
More tech → Read original →

Get the Tech newsletter

Curated tech stories, every morning. Free.

No spam. Unsubscribe anytime.