What it does
Build and serve large-language-model inference on NVIDIA GPUs
Not determinedOct 6, 2026702
View details OSS TAGS
Runs AI model inference and serves models through APIs or servers
Recently listed first
1–5 of 5 results
What it does
Build and serve large-language-model inference on NVIDIA GPUs
What it does
Serve large language and multimodal models with low-latency inference
What it does
bring text, image, audio, and video runtimes into one control plane
What it does
run GGUF models on a chosen mix of local CPUs and GPUs
What it does
one local runtime for acquiring models, running inference, and serving applications