Chinese Team Sets NanoGPT Speedrun Records via Architecture Redesign
Caiyun Technology leveraged dynamic attention and dense connections to push 124M-parameter training times down to 1 minute 16 seconds.

The Brief
Why it matters
China context
Editor's View
What to watch
- Whether mainstream open-source model frameworks adopt or adapt DCMHA and MUDD mechanisms for larger architectures.
- Follow-up attempts and counter-submissions by global research teams aiming to beat Caiyun Technology's 1-minute 16-second record on the NanoGPT Speedrun board.
- Further research measuring the engineering overhead and stability of dynamic routing when scaled to billion-parameter models.
Key Takeaways
- 1Caiyun Technology achieved a 1-minute 16-second training record on the NanoGPT Speedrun benchmark using eight H100 GPUs.
- 2The gains were driven by two novel architectures, DCMHA and MUDD, which introduce dynamic routing across attention heads and network layers.
- 3The code and custom kernels were accepted into the official open-source repository for public review and reproduction.
- 4The team noted that while the results validate dynamic data flow, adapting the methods to massive commercial models requires further engineering.
Sources
- 彩云科技:以底层架构创新两度改写开源AI竞速纪录--经济·科技--人民网 — People's Daily · 9/10/2026