Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
Original file line number Diff line number Diff line change
@@ -0,0 +1,47 @@
---
title: llama.cpp
icon: BookOpen
description: Configure llama.cpp as a custom endpoint in LibreChat.
---

[llama.cpp](https://github.com/ggml-org/llama.cpp) includes `llama-server`, which provides an OpenAI-compatible API for locally hosted models.

## Configuration

Set the same API key for LibreChat and `llama-server`. For example, add this to the `.env` file used by LibreChat:

```bash filename=".env"
LLAMA_CPP_API_KEY=replace-with-the-same-key
```

Start `llama-server` with a model alias and that key:

```bash
llama-server \
--model /path/to/model.gguf \
--alias local-model \
--port 8080 \
--api-key "replace-with-the-same-key"
```

Add this endpoint under `endpoints.custom` in your `librechat.yaml`:

```yaml filename="librechat.yaml"
- name: "llama.cpp"
apiKey: "${LLAMA_CPP_API_KEY}"
baseURL: "http://localhost:8080/v1"
models:
default: ["local-model"]
fetch: true
titleConvo: true
titleModel: "current_model"
summarize: false
summaryModel: "current_model"
modelDisplayLabel: "llama.cpp"
```

## Notes

- `--alias local-model` sets the model ID returned by the server, which must match the entry in `models.default`.
- `fetch: true` lets LibreChat discover models from the server's `/v1/models` endpoint.
- The server listens on `127.0.0.1:8080` by default. If LibreChat runs in Docker, use an address reachable from the API container instead of `localhost`. If you change the server's bind address, keep it on a trusted network and retain API-key protection.
Original file line number Diff line number Diff line change
Expand Up @@ -14,6 +14,7 @@
"helicone",
"huggingface",
"lemonade",
"llama-cpp",
"litellm",
"mistral",
"mlx",
Expand Down