TensorRT-LLM In-Flight Batching vs vLLM Continuous Batching
Three inference engines now use the same scheduling trick, but their real differences lie elsewhere.
Ingrid Zola
Section
1 story in Inference Runtime.
Three inference engines now use the same scheduling trick, but their real differences lie elsewhere.