Add mlops/llm-inference
This commit is contained in:
@@ -0,0 +1,43 @@
|
||||
---
|
||||
name: llm-inference
|
||||
description: "LLM inference: local GGUF inference (llama.cpp) and high-throughput serving (vLLM)."
|
||||
version: 2.0.0
|
||||
author: Hermes Agent
|
||||
license: MIT
|
||||
platforms: [linux, macos]
|
||||
metadata:
|
||||
hermes:
|
||||
tags: [llm, inference, llama-cpp, vllm, gguf, serving, deployment]
|
||||
related_skills: [huggingface-hub]
|
||||
---
|
||||
|
||||
# LLM Inference
|
||||
|
||||
Two inference approaches: local GGUF inference for development and vLLM for production serving.
|
||||
|
||||
## Section 1: llama.cpp — Local GGUF Inference
|
||||
|
||||
Run LLMs locally using llama.cpp with GGUF quantized models. Supports CPU, GPU, and mixed inference.
|
||||
|
||||
**When to use:** Local development, testing models, resource-constrained environments, offline inference.
|
||||
**Key features:** GGUF quantization, CPU/GPU/mixed, OpenAI-compatible server mode, HF Hub model discovery.
|
||||
**See:** `references/llama-cpp.md`.
|
||||
|
||||
## Section 2: vLLM — High-Throughput Serving
|
||||
|
||||
Deploy LLMs for production with vLLM. OpenAI API compatibility, PagedAttention, continuous batching, and quantization support.
|
||||
|
||||
**When to use:** Production LLM deployment, high-throughput serving, multi-model serving, needs OpenAI API compatibility.
|
||||
**Key features:** PagedAttention, continuous batching, tensor parallelism, GPTQ/AWQ quantization, OpenAI API server.
|
||||
**See:** `references/serving-llms-vllm.md`.
|
||||
|
||||
## Choosing Between Them
|
||||
|
||||
| Need | Tool |
|
||||
|------|------|
|
||||
| Local dev/testing | llama.cpp |
|
||||
| Resource-constrained | llama.cpp (with quantization) |
|
||||
| Production API serving | vLLM |
|
||||
| High throughput needed | vLLM |
|
||||
| OpenAI API compatibility | Both (via server modes) |
|
||||
| Multi-GPU serving | vLLM (tensor parallelism) |
|
||||
Reference in New Issue
Block a user