Morph Hiring Performance Engineer for Fastest Open Model Stack

Get the Tech newsletter
Daily tech — startups, AI labs, chips, the launches that shape the next decade. Free.
- Morph is hiring a Member of Technical Staff / performance engineer, with candidates expected to be top 1% across multiple parts of the inference stack.
- The role focuses on PD (prefill-decode) disaggregation research, plus autoresearch for kernels and custom speculative-decoding models.
- Morph's stack spans kernels, model serving, routing, autoscaling, and capacity — and serves the company's specialized code-generation models on custom inference infrastructure.
- Hires will work directly with the founders on problems determining how efficiently frontier-scale models can be served, on a small team with what the company describes as 'enormous compute' and immediate production impact.
- The hiring process includes a 2-day work trial, per Morph's listing.
Why it matters: Morph is staffing up at the inference layer — the segment where open-model economics get decided. By targeting PD disaggregation and custom speculative-decoding models, the company is betting that specialized code-generation serving can beat general-purpose providers on speed and cost, and is paying for that bet with senior-level talent recruited from the top of the inference stack.



