DSpark is a speculative-decoding framework described by researchers affiliated with Peking University and DeepSeek. It combines a semi-autoregressive draft model with confidence-based scheduling of verification work.
The research paper reports a 60–85% improvement in per-user generation speed over the MTP-1 production baseline at matched throughput in the DeepSeek-V4 serving system. This is a result under the paper's conditions, not a promise that every model, GPU or application gains the same speed or cost reduction.
The DeepSpec repository provides code for data preparation, draft-model training and evaluation, and links released research checkpoints. Availability of code does not make reproduction a one-click task or establish a universal reduction in API prices.
Date clarification: The arXiv record's first submission is July 6, 2026, after this site's June 27 archive date. It is used here as later verification; it is not presented as a paper that had already been published on the original date.
Editorial update, October 9, 2026: This archived brief has been checked against the verification sources below. The original publication date is retained; subsequent information is identified separately.
Verification sources: https://arxiv.org/abs/2607.05147 ; https://github.com/deepseek-ai/DeepSpec
