The system includes one or more processors configured to determine that a model is to be updated; determine a shard on which the model is to be deployed; determine whether to move the model to a different shard; in response to determining that the model is to be moved to the different shard, allocate the model to the different shard; and restart the different shard.