Efficiently serving machine-learned model computations with high throughput and low latency
An example method includes receiving input requests to process a plurality of input sequences using the machine-learned sequence processing model to generate a plurality of output sequences respectively corresponding to the plurality of input sequences; generating a plurality of initial attention tensors respectively for the plurality of input sequences, wherein: one or more respective initial attention tensors are generated for each respective input sequence in parallel over input elements of the respective input sequence; and the one or more respective initial attention tensors are generated in one or more batches having a first batch size using a prefill system that comprises one or more prefill computing devices and executes one or more layers of the machine-learned sequence processing model; and autoregressively generating, using the plurality of initial attention tensors, a plurality of output elements for each of the plurality of output sequences in one or more batches having a second batch size, wherein: the plurality of output elements are generated using a generation system that comprises one or more generation computing devices.