Back to News
news

KVCache over MoQT rev 01 lands on the IETF Datatracker

Large language model (LLM) inference involves two stages: prefill and decode. The prefill phase processes the prompt in parallel, generating the KVCache, which is then used by the decode phase to produce tokens sequentially. KVCache can be reused if the model and prompt is the same, reducing computing cost of the prefill. However, its large size makes efficient transfer challenging. Delivering these over architectures enabled by publish/subscribe transport like MoQT, allows local nodes to cache the KVCache to be later retrieved via new subscriptions, saving the bandwidth. This document specifies the transmission of KVCache over MoQT.

Source: IETF DatatrackerView source →

Draft Snapshot

  • Draft: draft-shi-moq-kvcache
  • Revision: 01
  • Last updated: 2025-10-11
  • Source: IETF Datatracker

Summary

Large language model (LLM) inference involves two stages: prefill and decode. The prefill phase processes the prompt in parallel, generating the KVCache, which is then used by the decode phase to produce tokens sequentially. KVCache can be reused if the model and prompt is the same, reducing computing cost of the prefill. However, its large size makes efficient transfer challenging. Delivering these over architectures enabled by publish/subscribe transport like MoQT, allows local nodes to cache the KVCache to be later retrieved via new subscriptions, saving the bandwidth. This document specifies the transmission of KVCache over MoQT.

Analysis

This draft is directly relevant to the MOQ ecosystem and worth tracking because it reflects current protocol work or adjacent implementation guidance from the IETF process.

Abstract

Large language model (LLM) inference involves two stages: prefill and decode. The prefill phase processes the prompt in parallel, generating the KVCache, which is then used by the decode phase to produce tokens sequentially. KVCache can be reused if the model and prompt is the same, reducing computing cost of the prefill. However, its large size makes efficient transfer challenging. Delivering these over architectures enabled by publish/subscribe transport like MoQT, allows local nodes to cache the KVCache to be later retrieved via new subscriptions, saving the bandwidth. This document specifies the transmission of KVCache over MoQT.

Stay ahead of MOQ

Get the latest IETF MOQ standards updates, protocol analysis, and ecosystem news delivered to your inbox.

Go deeper with MOQ Edge Pro

Weekly deep-dives, IETF standards tracking, and streaming tech trend reports — from $9/mo.

See plans →