mirror of
https://github.com/blakeblackshear/frigate.git
synced 2026-10-11 01:02:48 +03:00
Ollama image improvements (#24605)
CI / AMD64 Build (push) Canceled after 0s
CI / AMD64 Smoke Test (push) Canceled after 0s
CI / ARM Build (push) Canceled after 0s
CI / Jetson Jetpack 6 (push) Canceled after 0s
CI / AMD64 Extra Build (push) Canceled after 0s
CI / ARM Extra Build (push) Canceled after 0s
CI / Synaptics Build (push) Canceled after 0s
CI / Assemble and push default build (push) Canceled after 0s
CI / AMD64 Build (push) Canceled after 0s
CI / AMD64 Smoke Test (push) Canceled after 0s
CI / ARM Build (push) Canceled after 0s
CI / Jetson Jetpack 6 (push) Canceled after 0s
CI / AMD64 Extra Build (push) Canceled after 0s
CI / ARM Extra Build (push) Canceled after 0s
CI / Synaptics Build (push) Canceled after 0s
CI / Assemble and push default build (push) Canceled after 0s
* Support embeddinggemma for ollama * Refactor docs * Allow Ollama to probe cost of image tokens * Cleanup
This commit is contained in:
@@ -289,7 +289,7 @@ The only field that is valid at the camera level is `enabled`. In particular `mo
|
||||
|
||||
#### GenAI Provider
|
||||
|
||||
Frigate can send audio to a GenAI provider for transcription when that provider has the `transcribe` role. This is useful if you already run a GenAI provider, or if you do not have the CPU/GPU headroom for a local whisper model. Supported providers are **OpenAI**, **Azure OpenAI**, **Gemini**, and **llama.cpp** with an audio-capable model (a dedicated ASR model such as Qwen3-ASR, or a general multimodal model that accepts audio). Ollama is not supported as it has no audio input.
|
||||
Frigate can send audio to a GenAI provider for transcription when that provider has the `transcribe` role. This is useful if you already run a GenAI provider, or if you do not have the CPU/GPU headroom for a local whisper model. See [Provider support](/configuration/genai/genai_config#provider-support) for which providers can serve this role. The model must accept audio: either a dedicated ASR model such as Qwen3-ASR, or a general multimodal model that accepts audio.
|
||||
|
||||
To use a GenAI provider for audio transcription:
|
||||
|
||||
|
||||
@@ -43,7 +43,19 @@ genai:
|
||||
|
||||
The examples on this page all use `my_provider`, but the name is arbitrary and is only used to reference the provider elsewhere in the config (for example, `semantic_search.model`).
|
||||
|
||||
Each provider handles one or more **roles**: `chat`, `descriptions`, `embeddings`, and `transcribe`. A provider handles the first three by default; `transcribe` must always be listed explicitly, and is not available on Ollama, which has no audio input. Each role may be assigned to exactly one provider. Define a single provider if you want it to do everything, or split the roles across several providers using the `roles` option.
|
||||
Each provider handles one or more **roles**: `chat`, `descriptions`, `embeddings`, and `transcribe`. A provider handles the first three by default; `transcribe` must always be listed explicitly. Each role may be assigned to exactly one provider. Define a single provider if you want it to do everything, or split the roles across several providers using the `roles` option. Not every provider supports every role; see [Provider support](#provider-support).
|
||||
|
||||
### Provider support
|
||||
|
||||
| Provider | Descriptions | Chat | Embeddings | Transcription |
|
||||
| ----------------------------- | :----------: | :--: | :--------: | :-----------: |
|
||||
| llama.cpp (`llamacpp`) | ✅ | ✅ | ✅ | ✅ |
|
||||
| Ollama (`ollama`) | ✅ | ✅ | ✅ | ❌ |
|
||||
| OpenAI (`openai`) | ✅ | ✅ | ❌ | ✅ |
|
||||
| Azure OpenAI (`azure_openai`) | ✅ | ✅ | ❌ | ✅ |
|
||||
| Google Gemini (`gemini`) | ✅ | ✅ | ❌ | ✅ |
|
||||
|
||||
A ✅ means Frigate can use the provider for that feature. The configured model must also support it: a vision model for descriptions and chat, a multimodal embedding model for embeddings (see [Embedding models](#embedding-models)), and an audio-capable model for transcription. Some features also need extra provider setup, covered in each provider's section below. OpenAI-compatible servers use the `openai` provider, so they follow the OpenAI row.
|
||||
|
||||
If the provider you choose requires an API key, you may either directly paste it in your configuration, or store it in an environment variable prefixed with `FRIGATE_`.
|
||||
|
||||
@@ -73,9 +85,10 @@ You must use a vision-capable model with Frigate. The following models are recom
|
||||
|
||||
The `embeddings` role needs a different kind of model. Text queries are matched against the stored image embeddings, so the model must be trained to place images and text into the same vector space. A chat or description model will still return vectors when asked, but those vectors are not trained for retrieval and text searches will return poor matches with no error to indicate why.
|
||||
|
||||
| Model | Notes |
|
||||
| -------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
|
||||
| `qwen3-vl-embedding` | Multimodal embeddings for [Semantic Search](/configuration/semantic_search#genai-provider). Must be served by llama.cpp started with `--embeddings` and `--mmproj`. |
|
||||
| Model | Notes |
|
||||
| -------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
|
||||
| `embeddinggemma-2` | Multimodal embeddings for [Semantic Search](/configuration/semantic_search#genai-provider). Strong semantic search accuracy with efficient inference on a small model. |
|
||||
| `qwen3-vl-embedding` | Multimodal embeddings for [Semantic Search](/configuration/semantic_search#genai-provider). Good performance, large model that requires strong hardware for inference. |
|
||||
|
||||
#### Transcription models
|
||||
|
||||
@@ -152,6 +165,10 @@ genai:
|
||||
|
||||
Frigate queries the llama.cpp server for the model's context size at startup and logs it along with the other detected capabilities. If `context_size` is set in `provider_options`, that value is always used instead, even when the server reports its own.
|
||||
|
||||
#### Embeddings
|
||||
|
||||
To serve the `embeddings` role for [Semantic Search](/configuration/semantic_search#genai-provider), start the llama.cpp server with `--embeddings`, plus `--mmproj` for image support. See the [llama.cpp server documentation](https://github.com/ggml-org/llama.cpp/blob/master/tools/server/README.md) for details.
|
||||
|
||||
### Ollama
|
||||
|
||||
[Ollama](https://ollama.com/) allows you to self-host large language models and keep everything running locally. It is highly recommended to host this server on a machine with an Nvidia graphics card, or on a Apple silicon Mac for best performance.
|
||||
@@ -197,6 +214,10 @@ genai:
|
||||
</TabItem>
|
||||
</ConfigTabs>
|
||||
|
||||
#### Embeddings
|
||||
|
||||
Ollama can serve the `embeddings` role for [Semantic Search](/configuration/semantic_search#genai-provider). Embedding images requires Ollama 0.40.1 or newer and an embedding model with a vision encoder, such as `embeddinggemma-2:440m`. For a saved provider, the UI hides the role unless Ollama reports its model as an embedding model. Use a separate provider entry for the embedding model rather than adding the role to a vision chat model.
|
||||
|
||||
### OpenAI-Compatible
|
||||
|
||||
Frigate supports any provider that implements the OpenAI API standard. This includes self-hosted solutions like [vLLM](https://docs.vllm.ai/), [LocalAI](https://localai.io/), and other OpenAI-compatible servers.
|
||||
|
||||
@@ -133,13 +133,12 @@ Switching between V1 and V2 requires reindexing your embeddings. The embeddings
|
||||
|
||||
### GenAI Provider
|
||||
|
||||
Frigate can use a GenAI provider for semantic search embeddings when that provider has the `embeddings` role. Currently, only **llama.cpp** supports multimodal embeddings (both text and images).
|
||||
Frigate can use a GenAI provider for semantic search embeddings when that provider has the `embeddings` role. See [Provider support](/configuration/genai/genai_config#provider-support) for which providers can serve this role.
|
||||
|
||||
To use llama.cpp for semantic search:
|
||||
To use a GenAI provider for semantic search:
|
||||
|
||||
1. Configure a GenAI provider with `embeddings` in its `roles`.
|
||||
1. Configure a GenAI provider with `embeddings` in its `roles`, using a multimodal embedding model (both text and images). See [Embedding models](/configuration/genai/genai_config#embedding-models) for recommendations, and your provider's section of the [GenAI docs](/configuration/genai/genai_config) for any extra setup it needs.
|
||||
2. Set the semantic search model to the GenAI config key (e.g. `default`).
|
||||
3. Start the llama.cpp server with `--embeddings` and `--mmproj` for image support.
|
||||
|
||||
<ConfigTabs>
|
||||
<TabItem value="ui">
|
||||
@@ -174,8 +173,6 @@ semantic_search:
|
||||
</TabItem>
|
||||
</ConfigTabs>
|
||||
|
||||
The llama.cpp server must be started with `--embeddings` for the embeddings API, and a multi-modal embeddings model. See the [llama.cpp server documentation](https://github.com/ggml-org/llama.cpp/blob/master/tools/server/README.md) for details.
|
||||
|
||||
:::note
|
||||
|
||||
Switching between Jina models and a GenAI provider requires reindexing. Embeddings from different backends are incompatible.
|
||||
|
||||
Reference in New Issue
Block a user