TensorRT-LLM In-Flight Batching vs vLLM Continuous Batching
Three inference engines now use the same scheduling trick, but their real differences lie elsewhere.
Ingrid Zola
Staff Writer
Ingrid Zola is a staff writer at Runtime Review covering inference runtime. Based in Mumbai, Ingrid has written for Runtime Review since 2018.
1 story · Mumbai
Three inference engines now use the same scheduling trick, but their real differences lie elsewhere.