A Field Guide to Model-Based Speculative Drafting
This post compares trained drafter architectures for speculative decoding, from separate models to shared heads, feature predictors, and parallel drafters. It explains what each method changes and how it trades draft cost, acceptance length, reuse, and serving complexity.