vllm/triton.md at codex/add-pandas-and-datasets-to-requirements - vllm - Gitea: Git with a cup of tea

462 B

Raw Permalink Blame History

title
NVIDIA Triton

{ #deployment-triton }

The Triton Inference Server hosts a tutorial demonstrating how to quickly deploy a simple facebook/opt-125m model using vLLM. Please see Deploying a vLLM model in Triton for more details.