Compare commits

..
Author SHA1 Message Date
Josh HawkinsandGitHub e197b82804 fix save_attempts trimming (#24246) 2026-09-11 15:30:14 -05:00
Josh HawkinsandGitHub cd83fee1d6 Refactor Notices and System Health pane (#24243)
* refactor notices

* show startup message for enrichments in health pane

* tweaks
2026-09-11 11:53:55 -06:00
Nicolas MowenandGitHub afc335491d Revamp Face Recognition (#24236)
* Rename face model to face recognizer

* Refactor face detection into own module

* Return face and face landmarks

* Align faces with 5 points instead of just the eyes

* Add landmark validation to throw out images which do not fit a landmark

* Fix circular import and lock face detector for concurrent runs across threads

* Fix mypy
2026-09-10 10:58:26 -05:00
Josh HawkinsandGitHub 3175872a15 Report skipped detections in status bar (#24223)
* report skipped detections in the system notices pane

* revert notices and move to status bar

* shorten string
2026-09-08 15:45:55 -05:00
Nicolas MowenandGitHub 991af1370f Add ability to manually run Review Descriptions from the UI / API (#24222)
* Add ability to manually run Review Descriptions from the UI / API

* Fix not handling None type for call
2026-09-08 13:45:50 -05:00
Nicolas MowenandGitHub 35bec8499b Support main+sub stream exports (#24193)
* Support multi resolution exports

* Fix decoder text

* Add dropdown and ability to select export stream selection

* Fix for review comments

* Fix mypy

* Cleanup wording
2026-09-04 17:28:15 -05:00
Josh HawkinsandGitHub 45012d1c02 Improve System Health pane (#24188)
* build out system health pane

* tweaks

* fixes

* fix notice link so it opens the correct camera

* tweak language
2026-09-04 07:09:44 -06:00
Josh HawkinsandGitHub d014b62701 show disk space reclaimed by media sync (#24189) 2026-09-04 07:00:19 -06:00
Josh HawkinsandGitHub eef55045b6 Miscellaneous fixes (#24179)
* rename migration

* use ASC composite index for event camera and start_time

the planner still picks for newest-first single-camera queries but doesn't slow down full per-camera scans on a cold cache

* pause playback while the timeline range handles are up

* reduce multi camera seeded export range to 30m

allows both handlebars to fit within most desktop windows
2026-09-03 18:15:41 -05:00
2961449d4d Misc backend performance improvements (#23244)
* Add additional indicies on event and review tables.  Every events or timeline endpoint filters on event start time and camera, this should speed things up by avoiding a range scan on the table.

* Rewrite to use a CTE to leverage speedups by using sqllite internal optimization to do a single query instead of a starter query to get distinct labels and a subsequent loop of querys per distinct event labels.

Frigate is currently shipping sqlite 3.46.1, which is above the minimum version 3.25 needed for CTEs.

* Collapse a few sequential queries into a single one.

* Use peewee instead of rw sql for the CTE query.

* Slightly simplify review logic and avoid duplicating the json response for empty review IDs.

* Rerun ruff formatting.

* Remove 2x unnecessary index on reviewsegment, remove reference to prior code implementation in comment in event.py

* Editor fail, re-ruff format.

* Remove the CTE and restore the generator with sub-queries, which is more performance (thanks Nick and Blake for testing against your larger DB!)

* Update peewee index migration description

* Add on_conflict_ignore, replacing the try/catch/pass on IntegrityError

* Add a testcase for validating that on_conflict_ignore bypasses what was formerly an IntegrityError

* Change testcase to clarify that it covers the peewee behavior of on_conflict_ignore

---------

Co-authored-by: Greg <{ID}+{username}@users.noreply.github.com>
2026-09-03 15:53:59 -06:00
Josh HawkinsandGitHub 39f76e4a2c Add a notice registry and System Health tab (#24178)
* add a notice registry and System Health tab

Problems Frigate detects on its own (ffmpeg crash loops, stuck detectors, failed model downloads, recordings deleted before their retention period) only ever existed as log lines. This adds a `NoticeRegistry` in the main process backed by two tables, an `update_notice` IPC topic so producers in other processes can reach it through the dispatcher, an admin-only API and websocket topic, and a Health tab that lists them. Kinds declare their own mode, severity, and category in one place: state notices are resolved by their producer, event notices are dismissed by the user.

* treat a prerelease as behind its final release

* fix notices clearing early

* rename menu items and update docs

* don't resolve the update notice on a failed version lookup
2026-09-03 14:58:13 -06:00
Nicolas MowenandGitHub b6cf95a844 GenAI Chat Improvements (#24173)
* Initial tool approval implementation

* Cleanups and fixes

* Improve robustness of loading case
2026-09-03 09:02:22 -05:00
Nicolas MowenandGitHub 0453f945c7 Add options for review prompt style (#24166)
* Add script for testing genai review prompts

* Add option for prompt styling

* Add tests

* Update docs
2026-09-03 06:09:51 -06:00
Josh HawkinsandGitHub 141410f8e9 fix hardware stats crash when audio transcription is enabled (#24164) 2026-09-02 09:08:50 -06:00
Li XingyuandGitHub 1c3eebeebb Avoid killing recovered detector during restart (#24146) 2026-09-02 08:24:01 -06:00
Josh HawkinsandGitHub 505b976d19 fix latched loading spinner after cancelling a timeline selection (#24162)
isLoading and isBuffering only clear on playback progress, and scrubbing holds the player paused, so a source rebuild during a timeline selection left them set with nothing able to clear them. Cancelling made them visible as a spinner over an already-loaded frame, which stayed until the next manual seek. Leaving a scrub now clears them when the source is loaded and the element holds a frame.
2026-09-02 08:11:36 -06:00
Nicolas MowenandGitHub 5e2c59d27d Dynamically install and load detector dependencies (#24156)
* Dynamically install and load detector dependencies

* Cleanup

* Cleanup
2026-09-02 08:49:52 -05:00
Josh HawkinsandGitHub d97dd29561 Fix preview players at the hour rollover (#24157)
* fix preview players at the hour rollover

* clean up
2026-09-01 16:24:16 -05:00
Nicolas MowenandGitHub b2fc066e6c Refactor Hardware Stats (#24150)
* Refactor hardware stats to have consolidated ffmpeg, detector, and enrichments running.

* Cleanup hardware access that is not passed into the container

* Remove network stats from hardware refactor
2026-08-31 15:56:37 -06:00
Josh HawkinsandGitHub f0d7903965 add network isolation docs (#24148) 2026-08-31 11:57:05 -06:00
Josh HawkinsandGitHub 036786cb30 fix frozen time bounds on the recordings API (#24121)
`after` and `before` defaulted to `datetime.now()` in the function signature, so they were evaluated once at import. Requests that omitted them got a window ending at process start. They now resolve in the handler.
2026-08-31 11:56:05 -06:00
Josh HawkinsandGitHub 9876052cb4 Add a deny option for the proxy default role (#24145)
* backend

* frontend

* docs

* fix none default role casing and name reserved roles in the error

* reserve every casing of none as a role name
2026-08-31 08:34:15 -06:00
Josh HawkinsandGitHub 510e415a25 Container security hardening (phase 4) (#24140)
* Support read-only rootfs with self-signed certs in /config/tls

* Support read-only rootfs in s6 and pre-compile bytecode

* Assert read-only rootfs support in CI

* Document hardened read-only deployment

* keep certsync's cert selection identical to nginx's

* note the uid trade-off in user: mode

* fail fast when EXTRA_GROUPS or a missing media volume meets read_only

* keep nosuid and nodev on the /run tmpfs

* support read_only in the default mode

* don't take go2rtc down when the homekit file isn't writable

* lead with the hardware consequence of switching to user:

* refuse to write TLS material through a symlink as root

* note that memryx writes models to the root filesystem

* certsync watches whichever cert path nginx loaded
2026-08-31 05:41:32 -06:00
Nicolas MowenandGitHub ab20815584 Fix NPU turbo key and priviledges set (#24138) 2026-08-30 12:15:49 -05:00
Josh HawkinsandGitHub 890512b054 Container security hardening (phase 3, breaking) (#24081)
* Run the frigate service as the frigate user

* Run go2rtc as its own restricted user

* Run nginx as the frigate user with writable state in /tmp/nginx

* Disable bandwidth stats gracefully when not running as root

* Hand TensorRT model cache ownership to the runtime user

* Document non-root operation and per-hardware device access

* Create /media/frigate after the ownership sweep

* Assert non-root services, JWT migration, and escape hatch in CI

* only write the sweep sentinel when a media volume is mounted

* tolerate homekit config chown failures in the go2rtc run script

* chown the s6 log pipe so non-root nginx can reopen /dev/stdout

* set HOME to /config for non-root services

* run smoke nginx -t and the write probe as the runtime user

* re-own the nginx shm cache on service restart

* discard stdout for the unprivileged smoke nginx -t

* unwrap hard-wrapped prose in the installation docs

* report progress during the ownership sweep

* document EXTRA_GROUPS as the only device access path for dropped services

* expand the non-root device access docs with diagnosis steps and udev rules

* document network storage ownership and the remaining detector hardware

* skip lost+found during the ownership sweep

* hand /tmp/cache to the runtime user before services start

* make bundled models readable by the runtime user

* reload nginx by signaling the master instead of parsing its config as root

* harden root writes into unprivileged-owned paths

Restrict the sweep sentinel to a mount at or below /media/frigate so a
parent /media mount cannot bless a later-shadowed volume. Rebuild
/tmp/nginx root-owned each start so root's cp and tempio writes cannot
follow a symlink an unprivileged nginx planted in the previous run.

* collapse the duplicated sentinel comment

* add a service-runs-as-root helper for granular root services

* validate FRIGATE_ROOT_SERVICES and fail fast on unknown names

* let services listed in FRIGATE_ROOT_SERVICES skip the privilege drop

* record the root-services mode in the sentinel and sweep small trees each boot

* cache the runtime ids in the ownership helper

* chown recordings, previews, and exports to the runtime user at create

* chown the database files after init

* recommend FRIGATE_ROOT_SERVICES in the bandwidth stats warning

* assert granular root services in CI

* document FRIGATE_ROOT_SERVICES

* own every directory level created for a recording segment

* clear the cached runtime ids when ownership tests finish

* skip missing media paths in the per-boot ownership sweep

* clarify granular root services docs

* clean up

* install acl for device access grants

* grant runtime users access to mapped device nodes at boot

* assert device access grants in CI

* document automatic device access grants

* stop telling users device access needs host side setup

* clarify the non-root docs

* link the migration script to the repo

* group the manual device setup under one section

* harden against symlink attacks

/config is owned by the unprivileged runtime user after the ownership sweep, so root operations on files there could be redirected by a planted symlink.

- go2rtc HomeKit setup: replace the root yq/jq normalization and chown with an O_NOFOLLOW helper (prepare_homekit.py), so a symlink at go2rtc_homekit.yml can't redirect a root write or chown onto another file
- go2rtc binary override: ignore /config/go2rtc whenever the service runs as root, so a planted binary can't exec as root under FRIGATE_ROOT_SERVICES
- sweep sentinel: read and write it through safe-sentinel, which trusts only a root-owned regular file and never follows a symlink, so it can't be forged to skip the migration or symlinked to clobber a root file
- ownership sweep: chown with -execdir so a parent directory swapped for a symlink mid-walk can't redirect the chown out of the volume
- validate inputs: restrict DEVICE_ACL_PATHS to /dev, require nonzero numeric EXTRA_GROUPS, and reject PUID/PGID that collide with the go2rtc ids
- docs: correct the TLS key ownership note to match what actually happens

* tweak docs

* stop the ownership sweep chasing entries other mechanisms own

* keep custom binaries out of root services only under granular root
2026-08-30 08:14:43 -06:00
Josh HawkinsandGitHub a1fd978cab switch nginx-vod-module to the maintained dio-az fork (#24123)
The `v1.x` line is the same muxed fMP4 code as Kaltura's 1.31 with fixes backported, so the mapping JSON, manifest routes, and ffmpeg consumers are unchanged. The `MAX_CLIPS` patch applies as-is and the HEVC workaround is rewritten for the fork's reformatted source.

`vod_hls_version 6` is now explicit because the fork replaced Kaltura's automatic version calculation with a directive that defaults to 4 and only warns when fmp4 needs 6, so playlists were being stamped `EXT-X-VERSION:4` while carrying `EXT-X-MAP`. The `error_page 502 =404` hack is gone: https://github.com/kaltura/nginx-vod-module/issues/468 is a `vod_mode remote` bug and we're `mapped`, so those 502s were really `_vod_response` returning 404 upstream. The fork maps that through now, and the hack was also turning real 5xx into "no recordings".
2026-08-28 12:48:52 -06:00
Josh HawkinsandGitHub 5de5cee3c6 Fix LPR vehicle message for multiple models (#24119)
* use the camera's model for the lpr vehicle check

`FrigateConfig.model` became `models[]` in the detector refactor, so this didn't compile on 0.19. Each camera has its own detector/model pair now, so the check resolves the camera's scene with `getModelForCamera` instead of looking at every model.

* use camera config model

* fix export test
2026-08-28 12:26:05 -05:00
Josh Hawkins e99eb77ca1 Add Apple compatibility switch to the camera wizard (#24115)
* add apple compatibility switch to the camera wizard

* don't require every record stream to be h265
2026-08-27 20:30:36 -05:00
Josh Hawkins bb1e556ba9 Add onboarding wizard for new installations (#24102)
* add onboarding wizard for new users

* resolve hwaccel per camera and clarify recording retention

The hwaccel step listed every preset Frigate ships, so an Intel box was offered Raspberry Pi and Rockchip decoding, and the codec specific presets (`preset-intel-qsv-h264` vs `-h265`) were offered as global values that break as soon as two cameras use different codecs. `/hardware/hwaccel` now returns the decoding families the probed hardware can actually use, each carrying a preset per codec, and the wizard resolves the family against the detect stream codec the camera wizard already probed: one global `ffmpeg.hwaccel_args` when every camera agrees, per-camera `cameras.<name>.ffmpeg.hwaccel_args` when they don't. The global stays on `auto` in that case so cameras added later still resolve at startup. A gen13+ Intel machine keeps its QuickSync recommendation with mixed h264 and h265 cameras instead of dropping to vaapi.

The recording step's "Days to retain recordings" only wrote alert and detection retention, and the storage estimate under it assumed continuous recording. It now asks what to record in plain language, writes `record.continuous.days` to match, shows the estimate only for continuous, and drops the spinner arrows on the number input.

* clean up

* add light/dark mode icon switcher

* use yml as default config file extension when not found

* i18n tweaks

* gate the setup wizard on cameras instead of a config key

* render setup wizard steps by key

* share the setup wizard e2e helpers and mock users

* add an account step to the setup wizard

* add setup wizard account step e2e coverage

* cover the account step's restart behavior

* button consistency

* fix test

* docs

* fixes
2026-08-27 20:30:36 -05:00
Josh Hawkins 4f0c1b8ee7 Fix 500 when an event thumbnail file is empty (#24110)
* fix 500 when an event thumbnail file is empty

* fix test
2026-08-27 20:30:36 -05:00
Josh Hawkins a2085e27f6 Fix inconsistent export download filenames (#24111)
* fix inconsistent export download filenames

Zip entries in a case download were named from `Export.name`, the friendly display name, while an individual download uses the file name on disk. The two have always been formatted differently, so one export came out as `front_door_20260823_020615-20260823_020734_abc123.mp4` on its own and `front door 2026-08-23 020615 2026-08-23 020734.mp4` inside a zip. Zip entries now use the on-disk file name, and renaming an export renames its file, so there's only one name to download under. The rename is blocked while ffmpeg still holds the file.

* cap filename length and catch duplicate names

* fix export rename and stop blocking the event loop

* move the rename rollback off the event loop

* no awaits
2026-08-27 20:30:36 -05:00
Josh Hawkins 194e9bef69 Recording fixes (#24072)
* pin genai review frames to the main stream

* retain previews as long as either stream has recordings

* watch sub stream recording health separately from main

* reject record_sub on the same input as record and document the role

* derive recording paths from the cache segment timestamp

Recording paths carry one second of resolution, but since sub stream recording start times are resolved to fractional wall clock, anchored to the cache file mtime and chained to the previous segment's end. A stream cutting segments faster than once a second resolves consecutive segments into the same second, so two rows collide on the unique path index and the batch insert fails. The cache segment name is unique per camera stream and second by construction because ffmpeg names segments with strftime, so the recording path is now built from that timestamp while the row keeps the resolved start time. This also restores the path semantics from before sub stream recording, when start times came straight from the cache filename.

Nothing derives times from recording paths: playback offsets, stream switching, and export all use the row's start time, which is unchanged, and the recordings sync matches files by exact path string.

* keep the rest of a recording batch when one row conflicts

* only publish record_sub status when a sub stream is configured

* don't shadow camera_cfg when publishing empty cache streams

* back off restarts when a recording stream goes stale

* give the shared sub stream grace on any capture thread reset

* include segment details in recording discard warnings
2026-08-27 20:30:35 -05:00
Josh Hawkins 9d109bfd12 Container security hardening (phase 2) (#24068)
* Create frigate and go2rtc runtime users in the image

* Add single fix-ownership helper for volume permission migration

* Add init-usermod oneshot for PUID and PGID remapping

* Chown newly created runtime directories to the frigate user

* Run sentinel-guarded ownership sweep during prepare

* Add host-side volume permission migration script

* Guard log directory ownership for user-mode startup

* Fall back to plain s6-log when running without root

* Assert PUID remapping and sweep sentinel in CI smoke test

* Skip the ownership sweep in the devcontainer

* Pin FRIGATE_RUN_AS_ROOT in ownership tests

* Do not record the sweep as complete when a chown failed

* Validate PUID and PGID in the migration script

* Treat a failed ownership scan as an incomplete sweep

* Reject PUID and PGID of 0 during remapping

* Handle symlinks, dry runs, and sentinel write failures in the sweep

* Treat an absent sweep root as an incomplete sweep
2026-08-27 20:30:35 -05:00
Josh Hawkins dcda458a82 Container security hardening (phase 1) (#24061)
* Verify s6-overlay downloads against pinned checksums

* Verify go2rtc download against pinned checksums

The v1.9.14 release publishes no checksums file, just the bare per-platform binaries, so these digests come from a one-time fetch rather than upstream. That pins the artifact against later substitution, which is the realistic threat for a version we stay on for months, but it does not verify the original download. The stage moves from `ADD --link` to a script because `ADD --checksum` can't express an architecture-dependent URL.

* Verify main image downloads against pinned checksums

Covers everything the main image downloads on the default path: tempio, the hailort runtime tarball and wheel, the six ffmpeg builds, the libedgetpu deb, and the thirteen Intel driver debs. The hailort tarball was streamed straight into `tar`, which can't be verified before extraction, so it downloads to `/tmp` first. The three ffmpeg blocks per arch collapse into one `install_ffmpeg` helper since they only differed by URL and install dir, and the Intel debs go through a `fetch_intel_deb` helper for the same reason.

The Intel debs are the ones that mattered most here. They're installed as root with `dpkg` on the default amd64 path and had no verification at all. compute-runtime publishes a `ww<week>.sum` asset with every release and npu-driver published `checksum.sha256` on v1.19.0, so those eight digests came from upstream rather than from us. intel-graphics-compiler and level-zero publish none, so those five and everything else here come from a one-time fetch, which pins the artifact against later substitution but doesn't verify the original download. The comment above the map says which is which and how to refresh them, since npu-driver has stopped publishing sums since v1.19.0 and that provenance won't survive the next bump.

Still unpinned: `get-pip.py`, which is a rolling URL where a digest would just break the build on pypa's next edit, and the per-variant artifacts for Axera, Synaptics, and Jetson. apt repositories are out of scope since apt already verifies signatures.

* Restrict generated TLS key permissions

OpenSSL 3.x already writes the key at 600 on its own, so this pins the guarantee rather than fixing an observed leak: the mode no longer depends on the openssl version or the umask the service happens to start with. Only the generated pair is touched. User-mounted certs take the other branch and are never chmod'd, which matters when they're mounted read-only.

* Add security headers and server_tokens off

Adds `X-Content-Type-Options: nosniff` and `Referrer-Policy: strict-origin-when-cross-origin`, and turns off nginx version disclosure.

No `X-Frame-Options` and no CSP `frame-ancestors`. HA's Webpage card and iframe panels frame Frigate's own address cross-origin, and either header would break them silently with nothing in Frigate's logs to explain it. Ingress is same-origin and would survive `SAMEORIGIN`, but Frigate can't tell the two apart from inside the container. `security_headers.conf` is a plain file in the image rather than a generated one, so anyone who does want framing restrictions can bind-mount it.

`add_header` doesn't inherit into a block that declares its own, so the include goes in per block, all nine of them, including the four nested static-asset locations that serve the JS bundles. Those are the ones nosniff actually matters for.

The run script now reads `get_nginx_settings.py` once into a variable instead of shelling out per template. That script imports the frigate config machinery, which is noticeable on an SBC.

Not fixed here: `listen.conf` is included at server level and carries `Strict-Transport-Security`, so those same nine blocks already drop HSTS under TLS today. Folding it into this file would change existing TLS behavior on nine paths, so it needs its own PR.

* Restrict go2rtc config file permissions

* Log failed login attempts with source address

Failed logins returned a bare 401 and left nothing behind, so credential stuffing was invisible unless you were already watching nginx access logs. Both failure branches now log a warning with the attempted username and the client address.

The address comes from `get_remote_addr()`, the same helper the login rate limiter keys on, so the two agree on who the client is and the trusted-proxy handling is consistent. Logging a raw `x-forwarded-for` instead would let an attacker forge the source address in the very log line meant to catch them.

The response is unchanged and identical either way. Which factor failed is only visible in the log, never to the client, and the password is never logged.

* Recommend least-privilege container options in install docs

The compose generator pushed `privileged: true` into every file it produced, no matter what hardware you picked, and it's the default tab on the install page so it's what most people copy. It now emits `security_opt: no-new-privileges:true` instead, and only adds `privileged: true` for hardware that actually needs it, with the reason inline. MemryX is the only one today, since it needs to reach the max-manager. Rockchip and Synaptics only want privileged during initial setup and their documented end state is device mappings, so neither gets it.

`no-new-privileges` merges into the same `security_opt` block as any device-specific entries, so Rockchip still gets its `apparmor=unconfined` and `systempaths=unconfined` without a duplicate key.

The static example now has `privileged` commented out, and there's a short section on the options worth adding, with a note that `cap_drop: ALL` breaks `telemetry.stats.network_bandwidth` since nethogs needs NET_ADMIN/NET_RAW.

* Add amd64 container smoke test to CI

Boots the built amd64 image against a minimal config and asserts the two security headers, that the Server header no longer carries a version, that no frame-ancestors is present, that nginx accepts its own config, and the two file modes. This is also the harness the rest of the hardening work extends.

The two negative assertions are written as `if grep; then exit 1; fi` rather than `! grep`. Bash exempts a negated command from `set -e`, so the `!` form would have passed even with the version and frame-ancestors both present, which is the opposite of what a regression net is for.
2026-08-27 20:30:35 -05:00
Josh Hawkins 6c6683034e Tweaks (#24067)
* improve keyframes messages

* don't pad the labelmap with unknown

`load_labels()` prefilled 91 `unknown` entries before reading the label file, so any model with fewer than 91 classes kept that padding in `merged_labelmap` and `unknown` showed up as a selectable object type in the objects settings UI. The padding only existed so `RemoteObjectDetector.detect` could index the labelmap without a KeyError, and it didn't even cover the empty-file case or Frigate+, which never had a prefill. Both lookups now skip class ids the labelmap doesn't name and warn once per id.
2026-08-27 20:30:35 -05:00
Josh Hawkins 3ce3217db2 Add secrets.yaml and unify variable substitution sources (#24044)
* add secrets.yaml and merge substitution sources by precedence

FRIGATE_ENV_VARS was built once at import from container env and /run/secrets, and the environment_vars validator overwrote it unconditionally, so the block beat the deployment and nothing could be re-read. Sources are now separate dicts merged lowest to highest (environment_vars, secrets.yaml, container env, credentials directory), re-read at the top of every parse, and a collision warns once naming the winner. An undefined {FRIGATE_*} raises a ValueError subclass so pydantic reports the field instead of a KeyError traceback.

* use the shared substitution namespace in go2rtc config

The generator rebuilt the namespace itself from os.environ and a hardcoded /run/secrets, so it never saw environment_vars or CREDENTIALS_DIRECTORY, and str.format made any stray brace fatal. It now installs the FRIGATE_ names from environment_vars and substitutes streams the same way every other field does.

* read the exec override from an import time snapshot

environment_vars is exported into os.environ, and is_go2rtc_arbitrary_exec_allowed read os.environ live, so the config file could enable exec sources. Snapshot the variable at import, which runs before any config is loaded.

* docs

* clarify docs
2026-08-27 20:30:35 -05:00
Josh Hawkins 8a0c848914 add recognized plate picker to lpr known plates in settings (#24059) 2026-08-27 20:30:35 -05:00
Josh Hawkins fa4002cbe2 fix clip download deadlock from unread ffmpeg stderr (#24032)
ffmpeg's stderr was piped but never read, so recording segments that generate more than 64 KB of ffmpeg warnings blocked ffmpeg mid-write, stranding the streaming thread and its anyio threadpool token for good. Enough of those and every sync endpoint stops responding until restart. The trigger is how noisy the segments are, not how long the clip is.

Send stderr to a temp file instead, and guarantee ffmpeg teardown and playlist cleanup on every exit path, including client disconnect.

Also fixes two bugs the deadlock hid: the failure branch was unreachable because returncode is None mid-loop, so the playlist file leaked and ffmpeg's logs were never reported. Playlist files now get a unique name so concurrent requests for one range cannot delete each other's input.

Extracts the terminate helper motion search already had into frigate/util/ffmpeg.py, now shared by both streaming call sites.
2026-08-27 20:30:35 -05:00
Josh Hawkins a94b532655 fix the model lookup KeyError for cameras added at runtime (#24026) 2026-08-27 20:30:35 -05:00
Josh Hawkins fec73c887e Add import/export for camera group layouts and per-camera streaming settings (#24025)
* add import/export for camera group layouts and streaming settings

Camera group layouts and per-camera streaming settings are stored in the browser's IndexedDB, so they are tied to a single browser on a single device. Users with more than one device have to rebuild every group layout and re-pick every camera's stream settings by hand, and clearing browser data loses the work.

Add a Backup & Restore card to Settings > UI Settings that exports these settings to a JSON file and imports that file on another device. Import shows a confirmation dialog with per-section counts, switches for layouts, streaming settings, and UI preferences, and warnings about camera groups or cameras in the file that are not on this server.

Server-side storage is deliberately avoided. These are per-device presentation settings: a layout arranged for a desktop is wrong on a tablet, and continuous full-resolution streams that are free on a wired LAN are not on a phone. An explicit file moves settings only when the user chooses to move them.

Implementation notes:

- web/src/utils/uiSettingsTransfer.ts owns a registry of transferable IndexedDB keys. Each entry records whether the key is user-namespaced, matching which persistence hook wrote it, plus a zod schema for its value.
- Only registry-known keys are ever written, and only when their value passes that schema. The file format deliberately lets unknown keys survive parsing, so this filter is what prevents a hand-edited file from writing arbitrary storage keys or out-of-range values.
- Export falls back to the legacy un-namespaced key, because the username migration runs lazily on first mount of each owning hook.
- Streaming settings merge per group rather than replacing the whole map, so groups configured only on the receiving device survive.
- Import writes storage and then reloads, because useUserPersistence reads a key only on mount and StreamingSettingsProvider would otherwise write its stale in-memory state back over the import.
- playbackBandwidthEstimate, frigate-search-history, and live-layout are excluded: the first two are measurements and user data rather than preferences, and live-layout's default is derived from the device.

* merge imported streaming settings per camera instead of per group
2026-08-27 20:30:35 -05:00
Nicolas MowenandJosh Hawkins da135da0fb Implement UI for managing multiple models (#24023)
* Implement hardware detection and UI management

* Cleanup Frigate+ detection

* Don't count model as changed

* Fixes for audio map error

* Add descriptions

* Enforce that all model must exist

* Fix hardware picking

* Docs fixes

* WebUI cleanup

* Cleanup handling of scenes

* UI refinement

* Cleanup recommended UI

* test fixews
2026-08-27 20:30:35 -05:00
Josh Hawkins f5e398036e Base emergency cleanup on the streams a camera is currently recording (#24022)
* gate emergency cleanup bandwidth on the streams a camera currently records

* settle bandwidth samples per stream instead of per camera

* fix mypy
2026-08-27 20:30:35 -05:00
Nicolas MowenandJosh Hawkins 2dd700aa5a Refactor detector and model management (#23995)
* Refactor detector and model management

* Fix model resolution field
2026-08-27 20:30:35 -05:00
Ersa Oktavian RamadanandJosh Hawkins 378fbec416 Add audio labelmap grouping (#24004)
Allow audio classes to be grouped under a shared configured label.

Keep audio overrides separate from object labels and retain only the highest-scoring grouped detection.

Refs #23967
2026-08-27 20:30:35 -05:00
Josh Hawkins 91a93167d2 Show main and sub stream usage separately in Storage Metrics (#24015)
* backend

* frontend

* docs

* test

* report null instead of 0 for a stream with no cached bandwidth sample
2026-08-27 20:30:35 -05:00
Josh Hawkins dfe6428111 Refactor MQTT (#24010)
* refactor mqtt so that Frigate owns the transport lifecycle instead of delegating it to paho

* release the shutdown barrier on worker crash and replay retained publishes the broker never acked

* collapse in-flight retained values by topic and release the shutdown barrier from a finally

* replay the outage buffer before the publish queue so newer values are not reverted
2026-08-27 20:30:35 -05:00
Josh Hawkins 5e37b2c5c2 Refactor birdseye activity modes as a list and add alerts/detections (#24012)
* backend

* tests

* frontend and i18n

* e2e test schema

* docs
2026-08-27 20:30:35 -05:00
Josh Hawkins 36607133e5 Improve History's seek startup time and recordings query performance (#24011)
* serve a segment startup ladder so seeks begin playing sooner

nginx-vod was handed one 10s segment per recording file, so every playlist start had to download and decode a full segment before the first frame. Declare real keyframe data per clip and let nginx cut short leading segments from it.

- add vod_bootstrap_segment_durations 1000/2000/4000 so each playlist starts with 1s/2s/4s segments before settling at 10s
- emit real clip-relative keyFrameDurations (plus firstKeyFrameOffset when nonzero) from the recording keyframe index; rows without an index keep the whole-clip declaration, the only safe cut without keyframe knowledge
- drop the manifest's segment_duration field, which was always inert: nginx-vod parses only camelCase segmentDuration
- rebuild the player source at the seek target, quantized to a 10s grid, so the ladder applies to every seek and seek URLs stay repeatable for nginx's mapping and response caches
- route the seek model, in-range checks, and the stale-report guard through the source window rather than the chunk range
- bridge repositioning seeks (>2s from the last played timestamp) through the preview player and hold the release anchor one commit, so neither path paints a stale frame
- clear a pending loading timer before replacing it; an orphaned timer escaped onPlaying's clearTimeout and flashed loading mid-playback

* keep recordings queries on their indexes

Several recordings queries degraded into full scans or large sorts on big databases: the planner ignored index order, or the query shape gave it nothing tight to seek on. Reshape them into bounded seeks and add the composite index the per-stream lookups need.

- index recordings on (camera, stream_type, start_time DESC) and drop the (camera, stream_type) index it supersedes
- walk the recordings summary day by day with EXISTS probes and per-camera MIN/MAX seeks, skipping ahead over empty gaps instead of bucketing every row for the requested cameras
- run the summary endpoint on the event loop rather than the threadpool
- bound the unavailable-recordings query by start_time per camera and merge the results in Python
- bound the expire query's start_time so it seeks the retention window instead of scanning a camera's whole history
- enumerate deleted cameras with one index seek each rather than a camera NOT IN (...) scan
- compute bandwidth with segment_size filtered in a CASE projection; as a WHERE predicate it baited the planner into the (camera, segment_size) index plus a full sort of the camera's history
- fall back to a 1000-segment window when the recent 100 are all zero-size, so an ingest glitch doesn't report zero bandwidth
- limit the needs_refresh count instead of counting every segment
- cover sub-only and sparse calendar days, midnight-spanning day attribution, multi-camera gap merging, deleted-camera expiry, and zero-size segment runs

* fix mypy
2026-08-27 20:30:35 -05:00
Josh Hawkins 622fc97671 Enable PTZ control setup in the Add Camera Wizard (#23444)
* add ptz controls to camera via wizard when onvif has already been probed

* i18n

* add e2e test

* backend add and remove subscriber

* tweaks

* turn on switch by default if pan and/or tilt capability is available

* fix test
2026-08-27 20:30:35 -05:00
Josh Hawkins 5d807587ab Add sub stream recording with adaptive quality playback (#24009)
* add sub stream recording with adaptive quality playback

Optionally record a second, lower bitrate stream alongside the main
recording stream via a `record_sub` input role and `record.sub` config block, with its own retention windows.
Recordings rows now carry the stream type plus the media details needed to serve both streams from one manifest: video codec, audio presence, audio codec and rate, and a record-time keyframe index.

Playback resolves coverage across both streams and merges them into a single VOD sequence, falling back to a discontinuity manifest with per-clip init segments when the media signatures differ. The player exposes a quality selector, and an auto governor picks the stream from stall time, bandwidth, codec support, and the save-data hint.

* fix tests and i18n
2026-08-27 20:30:35 -05:00
Josh Hawkins 31bbf910c7 stop creating a config subscriber per capture thread (#24002) 2026-08-27 20:30:35 -05:00
Josh Hawkins f0d7c1d7d4 Guard lookups when adding/deleting cameras at runtime (#23994)
* Guard object processor queue handlers against unknown cameras

* Skip embeddings post processing for removed cameras

* End review segments for removed cameras

* Drop queued autotracker moves for removed cameras

* Release tracked event thumbnails when skipping a removed camera

* Add locked accessors for camera states

* Read camera states through the processor accessors

* Guard output and recording paths against cameras not yet known

* Resolve camera state once in ONVIF, notification, and transcription paths
2026-08-27 20:30:35 -05:00
Ersa Oktavian RamadanandJosh Hawkins a83219af56 Refactor Birdseye activity types as composable booleans (#23940)
* Add combined motion and object Birdseye mode

Add a motion_objects mode that keeps Birdseye active when motion is detected or a confirmed tracked object is present, including stationary objects.

Wire the mode through configuration, runtime commands, API schemas, documentation, and UI labels. Exclude false-positive trackers and add regression coverage for Birdseye activation and MQTT validation.

* Refactor Birdseye activity types as booleans

Replace combination-specific Birdseye modes with composable boolean activity types for motion, active objects, stationary objects, and continuous display.

Preserve legacy single-mode configuration and MQTT inputs, support canonical comma-separated MQTT combinations, and allow scalar YAML values to be replaced by nested settings through the config API.

* Preserve OpenVINO config translations

Regenerate the configuration translations with the OpenVINO detector schema available so the unrelated production detector labels remain intact.

* Preserve partial Birdseye mode overrides

Allow an empty activity selection with a canonical NONE MQTT state so partial camera and profile overrides can disable inherited flags without failing validation.

Add regression coverage for camera and profile inheritance, document the NONE contract, and keep the generated schema fixture scoped to Birdseye.

* Address Birdseye activity review feedback

Move scalar mode compatibility into the 0.18-1 config migration and reject empty activity selections instead of publishing a NONE state.

Pass activity signals through a frozen dataclass, preserve existing active-object tracker behavior, and require confirmed stationary objects. Revert the generic YAML mutation and cover migration, inheritance, MQTT, and activation regressions.

* Move Birdseye migration to 0.19

Use the 0.19-0 configuration revision for converting scalar Birdseye modes to composable activity flags, and update the migration regression coverage accordingly.

* Remove Birdseye migration test

Drop the dedicated config migration test as requested during review while retaining the 0.19-0 migration implementation.
2026-08-27 20:30:35 -05:00
Josh Hawkins 4a2fb2f09c Fix birdseye layout overlap with mixed landscape/portrait cameras (#22917)
* fix birdseye layout calculation

replace the two pass layout with a single pass pixel space algorithm

* add test
2026-08-27 20:30:35 -05:00
Nicolas MowenandJosh Hawkins 68893b28fc Don't require object type for parameter in categorized names tool 2026-08-27 20:30:35 -05:00
fbb904302c Dynamically resolve Intel NPU (#23761)
* Add support for newer Intel NPU busy time counter

* Resolve Intel NPU device dynamically

---------

Co-authored-by: Filious Louis <1417132+fjlouis@users.noreply.github.com>
2026-08-27 20:30:35 -05:00
DoFabienandJosh Hawkins e425ab5f90 Improve recording timeline and VOD query performance (#23862)
* Improve recording timeline and VOD query performance

* Add recording query boundary tests
2026-08-27 20:30:35 -05:00
Nicolas MowenandJosh Hawkins f486f7d57e GenAI Chat Prompt Refinements (#23864)
* Prompt refactoring and optimization

* Update spec
2026-08-27 20:30:35 -05:00
Nicolas MowenandJosh Hawkins c0adab1228 Update to 0.19 2026-08-27 20:30:35 -05:00
968 changed files with 31301 additions and 38982 deletions
+1
View File
@@ -55,6 +55,7 @@ Dahua
datasheet
debconf
deci
deepstack
defragment
devcontainer
DEVICEMAP
+2 -2
View File
@@ -26,8 +26,8 @@ body:
id: version
attributes:
label: Beta Version
description: Visible on the System Metrics page in the Web UI. Please include the full version including the build identifier (eg. 0.19.0-beta1, 0.19.0-8b72c7a, etc.)
placeholder: "0.19.0-beta1"
description: Visible on the System Metrics page in the Web UI. Please include the full version including the build identifier (eg. 0.18.0-beta1, 0.18.0-8b72c7a, etc.)
placeholder: "0.18.0-beta1"
validations:
required: true
- type: dropdown
+1 -3
View File
@@ -6,9 +6,7 @@ body:
value: |
Use this form to submit a reproducible bug in Frigate or Frigate's UI.
If you are running on Proxmox, please see the [Proxmox FAQ](https://github.com/blakeblackshear/frigate/discussions/23916) and reproduce the issue on a standard Docker install first (bare metal, or a VM running plain Debian/Ubuntu) before submitting here.
**⚠️ If you are running a beta version (0.19.0-beta or similar), please use the [Beta Support template](https://github.com/blakeblackshear/frigate/discussions/new?category=beta-support) instead.**
**⚠️ If you are running a beta version (0.18.0-beta or similar), please use the [Beta Support template](https://github.com/blakeblackshear/frigate/discussions/new?category=beta-support) instead.**
Before submitting your bug report, please ask the AI with the "Ask AI" button on the [official documentation site][ai] about your issue, [search the discussions][discussions], look at recent open and closed [pull requests][prs], read the [official Frigate documentation][docs], and read the [Frigate FAQ][faq] pinned at the Discussion page to see if your bug has already been fixed by the developers or reported by the community.
+15 -17
View File
@@ -23,7 +23,7 @@ jobs:
name: AMD64 Build
steps:
- name: Check out code
uses: actions/checkout@v7
uses: actions/checkout@v6
with:
persist-credentials: false
- name: Set up QEMU and Buildx
@@ -49,7 +49,7 @@ jobs:
- amd64_build
steps:
- name: Check out code
uses: actions/checkout@v7
uses: actions/checkout@v6
with:
persist-credentials: false
- name: Set up QEMU and Buildx
@@ -144,11 +144,9 @@ jobs:
done'
- name: Assert device access grants
run: |
# a fake accelerator node created after boot, then the oneshot re-run.
# /command is on PATH only for s6-supervised services, and the
# with-contenv shebang resolves its execline helpers through PATH
# a fake accelerator node created after boot, then the oneshot re-run
docker exec frigate mknod /dev/apex_9 c 120 99
docker exec frigate sh -c 'export PATH=/command:$PATH; exec /etc/s6-overlay/s6-rc.d/init-devices/run'
docker exec frigate /etc/s6-overlay/s6-rc.d/init-devices/run
acl=$(docker exec frigate getfacl -p /dev/apex_9)
echo "$acl"
echo "$acl" | grep -q "user:frigate:rw-"
@@ -156,7 +154,7 @@ jobs:
# the usb tree gets recursive grants plus a default ACL that
# newly created nodes inherit (the Coral re-enumeration path)
docker exec frigate sh -c 'mkdir -p /dev/bus/usb/001 && mknod /dev/bus/usb/001/002 c 189 1'
docker exec frigate sh -c 'export PATH=/command:$PATH; exec /etc/s6-overlay/s6-rc.d/init-devices/run'
docker exec frigate /etc/s6-overlay/s6-rc.d/init-devices/run
docker exec frigate getfacl -p /dev/bus/usb/001 | grep -q "user:frigate:rwx"
docker exec frigate sh -c 'mknod /dev/bus/usb/001/099 c 189 98 && chmod 664 /dev/bus/usb/001/099'
inherited=$(docker exec frigate getfacl -p /dev/bus/usb/001/099)
@@ -173,7 +171,7 @@ jobs:
# hardware that is absent must stay silent: the literal table entries
# are not globs, so nullglob does not drop them and only an existence
# check keeps them from warning on every boot
out=$(docker exec frigate sh -c 'export PATH=/command:$PATH; exec /etc/s6-overlay/s6-rc.d/init-devices/run')
out=$(docker exec frigate /etc/s6-overlay/s6-rc.d/init-devices/run)
echo "$out"
if echo "$out" | grep -q "WARN"; then
echo "grant warned about device nodes that do not exist"; exit 1
@@ -303,7 +301,7 @@ jobs:
# /run must allow exec: S6_READ_ONLY_ROOT has s6 copy its service
# scripts there and run them, and --tmpfs defaults to noexec
docker run -d --name frigate-ro --shm-size 256m \
--read-only --tmpfs /tmp:rw,size=1g --tmpfs /run:exec,nosuid,nodev,mode=0755,uid=1000,gid=1000 \
--read-only --tmpfs /tmp:rw,size=1g --tmpfs /run:exec,nosuid,nodev,mode=0755 \
--user 1000:1000 \
--security-opt no-new-privileges:true \
-v /tmp/frigate-config-ro:/config \
@@ -397,7 +395,7 @@ jobs:
# setfacl under a read-only rootfs, which nothing else covers:
# init-devices exits early under --user, so that path is never reached
docker exec frigate-rod mknod /dev/apex_9 c 120 99
docker exec frigate-rod sh -c 'export PATH=/command:$PATH; exec /etc/s6-overlay/s6-rc.d/init-devices/run'
docker exec frigate-rod /etc/s6-overlay/s6-rc.d/init-devices/run
docker exec frigate-rod getfacl -p /dev/apex_9 | grep -q "user:frigate:rw-"
docker rm -f frigate-rod
- name: "Assert switching that install to user: still starts"
@@ -405,7 +403,7 @@ jobs:
# the config dir above now holds a go2rtc-owned go2rtc_homekit.yml,
# which user: keeps readable but not writable (no supplementary groups)
docker run -d --name frigate-rod-user --shm-size 256m \
--read-only --tmpfs /tmp:rw,size=1g --tmpfs /run:exec,nosuid,nodev,mode=0755,uid=1000,gid=1000 \
--read-only --tmpfs /tmp:rw,size=1g --tmpfs /run:exec,nosuid,nodev,mode=0755 \
--user 1000:1000 \
-v /tmp/frigate-config-rod:/config \
-v /tmp/frigate-media-rod:/media/frigate \
@@ -426,7 +424,7 @@ jobs:
name: ARM Build
steps:
- name: Check out code
uses: actions/checkout@v7
uses: actions/checkout@v6
with:
persist-credentials: false
- name: Set up QEMU and Buildx
@@ -461,7 +459,7 @@ jobs:
name: Jetson Jetpack 6
steps:
- name: Check out code
uses: actions/checkout@v7
uses: actions/checkout@v6
with:
persist-credentials: false
- name: Set up QEMU and Buildx
@@ -492,7 +490,7 @@ jobs:
- amd64_build
steps:
- name: Check out code
uses: actions/checkout@v7
uses: actions/checkout@v6
with:
persist-credentials: false
- name: Set up QEMU and Buildx
@@ -533,7 +531,7 @@ jobs:
- arm64_build
steps:
- name: Check out code
uses: actions/checkout@v7
uses: actions/checkout@v6
with:
persist-credentials: false
- name: Set up QEMU and Buildx
@@ -558,7 +556,7 @@ jobs:
- arm64_build
steps:
- name: Check out code
uses: actions/checkout@v7
uses: actions/checkout@v6
with:
persist-credentials: false
- name: Set up QEMU and Buildx
@@ -590,7 +588,7 @@ jobs:
with:
string: ${{ github.repository }}
- name: Log in to the Container registry
uses: docker/login-action@dbcb813823bdd20940b903addbd779551569679f
uses: docker/login-action@184bdaa0721073962dff0199f1fb9940f07167d1
with:
registry: ghcr.io
username: ${{ github.actor }}
+13 -10
View File
@@ -16,10 +16,10 @@ jobs:
name: Web - Lint
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v7
- uses: actions/checkout@v6
with:
persist-credentials: false
- uses: actions/setup-node@v7
- uses: actions/setup-node@v6
with:
node-version: 20.x
- run: npm install
@@ -35,10 +35,10 @@ jobs:
name: Web - Test
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v7
- uses: actions/checkout@v6
with:
persist-credentials: false
- uses: actions/setup-node@v7
- uses: actions/setup-node@v6
with:
node-version: 20.x
- run: npm install
@@ -46,15 +46,18 @@ jobs:
- name: Build web
run: npm run build
working-directory: ./web
# - name: Test
# run: npm run test
# working-directory: ./web
web_e2e:
name: Web - E2E Tests
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v7
- uses: actions/checkout@v6
with:
persist-credentials: false
- uses: actions/setup-node@v7
- uses: actions/setup-node@v6
with:
node-version: 20.x
- run: npm install
@@ -83,11 +86,11 @@ jobs:
name: Python Checks
steps:
- name: Check out the repository
uses: actions/checkout@v7
uses: actions/checkout@v6
with:
persist-credentials: false
- name: Set up Python ${{ env.DEFAULT_PYTHON }}
uses: actions/setup-python@v7.0.0
uses: actions/setup-python@v5.4.0
with:
python-version: ${{ env.DEFAULT_PYTHON }}
- name: Install requirements
@@ -106,10 +109,10 @@ jobs:
name: Python Tests
steps:
- name: Check out code
uses: actions/checkout@v7
uses: actions/checkout@v6
with:
persist-credentials: false
- uses: actions/setup-node@v7
- uses: actions/setup-node@v6
with:
node-version: 20.x
- name: Install devcontainer cli
+2 -2
View File
@@ -10,7 +10,7 @@ jobs:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v7
- uses: actions/checkout@v6
with:
persist-credentials: false
- id: lowercaseRepo
@@ -18,7 +18,7 @@ jobs:
with:
string: ${{ github.repository }}
- name: Log in to the Container registry
uses: docker/login-action@dbcb813823bdd20940b903addbd779551569679f
uses: docker/login-action@184bdaa0721073962dff0199f1fb9940f07167d1
with:
registry: ghcr.io
username: ${{ github.actor }}
+1 -1
View File
@@ -160,7 +160,7 @@ When reviewing code, do NOT comment on:
### Code Quality
- **Linting**: ESLint (see `web/eslint.config.js`)
- **Linting**: ESLint (see `web/.eslintrc.cjs`)
- **Formatting**: Prettier with Tailwind CSS plugin
- **Type Safety**: TypeScript strict mode enabled
+1 -1
View File
@@ -1,4 +1,4 @@
ruff == 0.15.20
# types
types-peewee == 4.0.*
types-peewee == 3.17.*
+17 -17
View File
@@ -1,17 +1,17 @@
aiofiles == 25.1.*
click == 8.5.*
aiofiles == 24.1.*
click == 8.1.*
# FastAPI
aiohttp == 3.12.*
starlette == 0.47.*
starlette-context == 0.5.*
starlette-context == 0.4.*
fastapi[standard-no-fastapi-cloud-cli] == 0.116.*
uvicorn == 0.52.*
uvicorn == 0.35.*
slowapi == 0.1.*
joserfc == 1.6.*
cryptography == 46.0.*
joserfc == 1.2.*
cryptography == 44.0.*
pathvalidate == 3.3.*
markupsafe == 3.0.*
python-multipart == 0.0.31
python-multipart == 0.0.26
# Classification Model Training
tensorflow == 2.19.* ; platform_machine == 'aarch64'
tensorflow-cpu == 2.19.* ; platform_machine == 'x86_64'
@@ -26,15 +26,15 @@ psutil == 7.1.*
pydantic == 2.10.*
git+https://github.com/fbcotter/py3nvml#egg=py3nvml
pytz == 2025.*
pyzmq == 27.1.*
pyzmq == 26.2.*
ruamel.yaml == 0.18.*
tzlocal == 5.2
requests == 2.33.*
requests == 2.32.*
types-requests == 2.32.*
norfair == 2.3.*
setproctitle == 1.3.*
ws4py == 0.5.*
unidecode == 1.4.*
unidecode == 1.3.*
titlecase == 2.4.*
# Image Manipulation
numpy == 1.26.*
@@ -51,26 +51,26 @@ google-genai == 1.58.*
ollama == 0.6.*
openai == 1.65.*
# push notifications
py-vapid == 1.9.4
py-vapid == 1.9.*
pywebpush == 2.0.*
# alpr
pyclipper == 1.4.*
pyclipper == 1.3.*
shapely == 2.0.*
rapidfuzz==3.12.*
# HailoRT
argcomplete==2.0.*
contextlib2==0.6.*
future==0.18.*
netaddr==1.3.*
netaddr==0.8.*
netifaces==0.10.*
prometheus-client == 0.26.*
prometheus-client == 0.21.*
# TFLite
tflite_runtime @ https://github.com/frigate-nvr/TFlite-builds/releases/download/v2.17.1/tflite_runtime-2.17.1-cp311-cp311-linux_x86_64.whl; platform_machine == 'x86_64'
tflite_runtime @ https://github.com/feranick/TFlite-builds/releases/download/v2.17.1/tflite_runtime-2.17.1-cp311-cp311-linux_aarch64.whl; platform_machine == 'aarch64'
# audio transcription
sherpa-onnx==1.13.*
faster-whisper==1.2.*
sherpa-onnx==1.12.*
faster-whisper==1.1.*
librosa==0.11.*
soundfile==0.13.*
# Memory profiling
memray == 1.20.*
memray == 1.15.*
@@ -166,9 +166,7 @@ http {
include auth_request.conf;
types {
video/mp4 mp4;
image/jpeg jpg jpeg;
image/png png;
image/webp webp;
image/jpeg jpg;
}
expires 7d;
@@ -343,6 +341,13 @@ http {
add_header Cache-Control "public";
}
location /fonts/ {
access_log off;
expires 1y;
include security_headers.conf;
add_header Cache-Control "public";
}
location /locales/ {
access_log off;
include security_headers.conf;
@@ -369,7 +374,7 @@ http {
sub_filter '"/BASE_PATH/assets/' '"$http_x_ingress_path/assets/';
sub_filter '"/BASE_PATH/locales/' '"$http_x_ingress_path/locales/';
sub_filter '"/BASE_PATH/monacoeditorwork/' '"$http_x_ingress_path/assets/';
sub_filter 'return`/BASE_PATH/`' 'return window.baseUrl';
sub_filter 'return"/BASE_PATH/"' 'return window.baseUrl';
sub_filter '<body>' '<body><script>window.baseUrl="$http_x_ingress_path/";</script>';
sub_filter_types text/css application/javascript;
sub_filter_once off;
+1
View File
@@ -1,5 +1,6 @@
# Nvidia ONNX Runtime GPU Support
--extra-index-url 'https://pypi.nvidia.com'
cython==3.0.*; platform_machine == 'x86_64'
nvidia-cuda-cupti-cu12==12.8.90; platform_machine == 'x86_64'
nvidia-cublas-cu12==12.8.4.1; platform_machine == 'x86_64'
nvidia-cudnn-cu12==9.8.0.87; platform_machine == 'x86_64'
+24 -6
View File
@@ -36,13 +36,13 @@ edgeTPU:
height: 320 # <--- should match the imgsize of the model, typically 320
path: /config/model_cache/yolov9-s-relu6-best_320_int8_edgetpu.tflite
labelmap_path: /config/labels-coco17.txt
hailo:
title: Hailo
hailo8l:
title: Hailo-8/Hailo-8L
models:
- key: yolo
label: YOLO
recommended: true
download: If no custom model path or URL is provided, the Hailo detector automatically downloads the default model (YOLOv6n) from the Hailo Model Zoo on first startup, choosing the build that matches the attached device. Once cached under `/config/model_cache`, the model works fully offline.
download: If no custom model path or URL is provided, the Hailo detector automatically downloads the default model (YOLOv6n) from the Hailo Model Zoo on first startup based on the detected hardware. Once cached under `/config/model_cache/hailo`, the model works fully offline.
ui: |-
Navigate to **Settings > System > Detection models** and select **Hailo** from the **Hardware** dropdown. Then, on the same model, open the **Custom Model** tab and configure the model settings:
@@ -60,7 +60,7 @@ hailo:
yaml: |-
models:
- devices:
- hailo:PCIe
- hailo8l:PCIe
width: 320
height: 320
input_tensor: nhwc
@@ -101,7 +101,7 @@ hailo:
yaml: |-
models:
- devices:
- hailo:PCIe
- hailo8l:PCIe
width: 300
height: 300
input_tensor: nhwc
@@ -824,6 +824,24 @@ cpu:
models:
- devices:
- cpu:3
deepstack:
title: DeepStack / CodeProject.AI
models:
- key: yolo
label: YOLO
recommended: true
download: This detector runs object detection over the network against a CodeProject.AI or DeepStack server, so no model is downloaded into Frigate itself. Visit the [CodeProject.AI official website](https://www.codeproject.com/Articles/5322557/CodeProject-AI-Server-AI-the-easy-way) to download and install the AI server on your preferred device (e.g. Raspberry Pi, Nvidia Jetson, or other compatible hardware) before configuring the detector.
ui: |-
Navigate to **Settings > System > Detection models** and add a model. The CodeProject.AI server is not reported by the hardware probe, so set `devices` to `deepstack:http://<your_codeproject_ai_server_ip>:<port>/v1/vision/detection` in YAML.
| Field | Value |
| ------------- | ---------------------------------------------------------------------- |
| **API URL** | `http://<your_codeproject_ai_server_ip>:<port>/v1/vision/detection` |
| **API Timeout** | `0.1` (seconds) |
yaml: |-
models:
- devices:
- deepstack:http://<your_codeproject_ai_server_ip>:<port>/v1/vision/detection
memryx:
title: MemryX
models:
@@ -1015,7 +1033,7 @@ synaptics:
- key: ssd
label: SSD MobileNet
recommended: true
download: A synap model is provided in the container at `/synaptics/mobilenet.synap` and is used by this detector type by default. The model comes from the [Synap-release Github](https://github.com/synaptics-astra/synap-release/tree/v1.5.0/models/dolphin/object_detection/coco/model/mobilenet224_full80).
download: A synap model is provided in the container at `/mobilenet.synap` and is used by this detector type by default. The model comes from the [Synap-release Github](https://github.com/synaptics-astra/synap-release/tree/v1.5.0/models/dolphin/object_detection/coco/model/mobilenet224_full80).
ui: |-
Navigate to **Settings > System > Detection models** and select **Synaptics NPU** from the **Hardware** dropdown. Then, on the same model, open the **Custom Model** tab and configure:
+4 -12
View File
@@ -800,7 +800,7 @@ lpr:
# to Google or OpenAI's LLMs to generate descriptions. GenAI features can be configured at
# the camera level to enhance privacy for indoor cameras.
# NOTE: genai is a map of named providers. Each key is a name you choose for the provider,
# and each role (chat, descriptions, embeddings, transcribe) may be assigned to exactly one provider.
# and each role (chat, descriptions, embeddings) may be assigned to exactly one provider.
genai:
# Required: name of the provider (chosen by you, used to reference it elsewhere)
my_provider:
@@ -813,13 +813,11 @@ genai:
# Required: The model to use with the provider.
model: gemini-1.5-flash
# Optional: Roles this provider handles (default: shown below)
# Each role (chat, descriptions, embeddings, transcribe) must be assigned to exactly
# one provider.
# Each role (chat, descriptions, embeddings) must be assigned to exactly one provider.
roles:
- chat
- descriptions
- embeddings
- transcribe
# Optional additional args to pass to the GenAI Provider (default: None)
provider_options:
keep_alive: -1
@@ -832,19 +830,13 @@ genai:
audio_transcription:
# Optional: Enable live and speech event audio transcription (default: shown below)
enabled: False
# Optional: The transcription backend (default: shown below)
# Either 'whisper' for Frigate's built-in local models, or the name of a genai
# provider that has 'transcribe' in its roles. device and model_size are ignored
# when a genai provider is named.
model: whisper
# Optional: The device to run the models on for live transcription. (default: shown below)
device: CPU
# Optional: Set the model size used for live transcription. (default: shown below)
model_size: small
# Optional: Set the language used for transcription translation. (default: shown below)
# Use 'auto' to let the model detect the language, or a language code from
# https://github.com/openai/whisper/blob/main/whisper/tokenizer.py#L10
language: auto
# List of language codes: https://github.com/openai/whisper/blob/main/whisper/tokenizer.py#L10
language: en
# Optional: Configuration for classification models
classification:
+7 -79
View File
@@ -204,7 +204,7 @@ Frequently-heard labels like `speech` can generate a lot of events, and each eve
### Audio Transcription
Frigate supports fully local audio transcription using either `sherpa-onnx` or OpenAI's open-source Whisper models via `faster-whisper`, and can alternatively offload transcription to a [GenAI provider](#genai-provider). The goal of this feature is to support Semantic Search for `speech` audio events. Frigate is not intended to act as a continuous, fully-automatic speech transcription service. Automatically transcribing all speech (or queuing many audio events for transcription) requires substantial CPU (or GPU) resources and is impractical on most systems. For this reason, transcriptions for events are initiated manually from the UI or the API rather than being run continuously in the background.
Frigate supports fully local audio transcription using either `sherpa-onnx` or OpenAI's open-source Whisper models via `faster-whisper`. The goal of this feature is to support Semantic Search for `speech` audio events. Frigate is not intended to act as a continuous, fully-automatic speech transcription service. Automatically transcribing all speech (or queuing many audio events for transcription) requires substantial CPU (or GPU) resources and is impractical on most systems. For this reason, transcriptions for events are initiated manually from the UI or the API rather than being run continuously in the background.
:::info
@@ -224,7 +224,6 @@ To enable transcription, configure it globally and optionally disable for specif
**Global:** Navigate to <NavPath path="Settings > Enrichments > Audio transcription" />.
- Set **Enable audio transcription** to on
- Set **Audio transcription model or GenAI provider name** to `whisper` for Frigate's built-in local models, or to the name of a GenAI provider
- Set **Transcription device** to the desired device
- Set **Model size** to the desired size
@@ -236,7 +235,6 @@ To enable transcription, configure it globally and optionally disable for specif
```yaml
audio_transcription:
enabled: True
model: whisper
device: ...
model_size: ...
```
@@ -265,88 +263,20 @@ The optional config parameters that can be set at the global level include:
- **`enabled`**: Enable or disable the audio transcription feature.
- Default: `False`
- It is recommended to only configure the features at the global level, and enable it at the individual camera level.
- **`model`**: The transcription backend.
- Default: `whisper`
- `whisper` uses Frigate's built-in local models, described by `device` and `model_size` below.
- Any other value must name a key in your `genai` config whose entry has `transcribe` in its `roles`. See [GenAI Provider](#genai-provider).
- **`device`**: Device to use to run transcription and translation models.
- Default: `CPU`
- This can be `CPU` or `GPU`. The `sherpa-onnx` models are lightweight and run on the CPU only. The `whisper` models can run on GPU but are only supported on CUDA hardware.
- Ignored when `model` names a GenAI provider.
- **`model_size`**: The size of the model used for live transcription.
- Default: `small`
- This can be `small` or `large`. The `small` setting uses `sherpa-onnx` models that are fast, lightweight, and always run on the CPU but are not as accurate as the `whisper` model.
- This config option applies to **live transcription only**. With `model: whisper`, recorded `speech` events always use a different `whisper` model (and can be accelerated for CUDA hardware if available with `device: GPU`).
- Ignored when `model` names a GenAI provider.
- **`language`**: Defines the language used to transcribe and translate `speech` audio events (and live audio only if using the `large` model or a GenAI provider).
- Default: `auto`
- `auto` lets the model detect the language itself, which most models do well. Set an explicit language only if detection is picking the wrong one.
- Otherwise you must use a valid [language code](https://github.com/openai/whisper/blob/main/whisper/tokenizer.py#L10).
- This config option applies to **live transcription only**. Recorded `speech` events will always use a different `whisper` model (and can be accelerated for CUDA hardware if available with `device: GPU`).
- **`language`**: Defines the language used by `whisper` to translate `speech` audio events (and live audio only if using the `large` model).
- Default: `en`
- You must use a valid [language code](https://github.com/openai/whisper/blob/main/whisper/tokenizer.py#L10).
- Transcriptions for `speech` events are translated.
- Live audio is translated only if you are using the `large` model. The `small` `sherpa-onnx` model is English-only.
The only field that is valid at the camera level is `enabled`. In particular `model` is global only: the transcription backend is a process-wide resource shared by every camera.
#### GenAI Provider
Frigate can send audio to a GenAI provider for transcription when that provider has the `transcribe` role. This is useful if you already run a GenAI provider, or if you do not have the CPU/GPU headroom for a local whisper model. Supported providers are **OpenAI**, **Azure OpenAI**, **Gemini**, and **llama.cpp** with an audio-capable model (a dedicated ASR model such as Qwen3-ASR, or a general multimodal model that accepts audio). Ollama is not supported as it has no audio input.
To use a GenAI provider for audio transcription:
1. Configure a GenAI provider with `transcribe` in its `roles`.
2. Set the audio transcription model to that GenAI config key (e.g. `whisper_cloud`).
<ConfigTabs>
<TabItem value="ui">
Navigate to <NavPath path="Settings > Enrichments > Audio transcription" />.
| Field | Description |
| ---------------------------------------------------- | ---------------------------------------------------------------------------------------------------------- |
| **Audio transcription model or GenAI provider name** | Set to the GenAI config key (e.g. `whisper_cloud`) to use a configured GenAI provider for transcription |
The GenAI provider must also be configured with the `transcribe` role under <NavPath path="Settings > Enrichments > Generative AI" />.
</TabItem>
<TabItem value="yaml">
```yaml
genai:
whisper_cloud:
provider: openai
api_key: your-api-key
model: gpt-transcribe
roles:
- transcribe
audio_transcription:
enabled: True
model: whisper_cloud
language: en
```
</TabItem>
</ConfigTabs>
:::warning
**Give `transcribe` its own `genai` entry.** A `genai` entry has a single `model` string that is shared by every role it holds, so `roles: [descriptions, transcribe]` would send the same model name to both the chat endpoint and the transcription endpoint. Transcription models and chat models are almost never the same model, so define a dedicated entry as shown above.
:::
:::warning
**Live transcription against a metered provider is billed continuously.** In live mode Frigate uploads an overlapping ~2 second window of audio roughly once per second, per camera, for as long as audio stays above that camera's `audio.min_volume`. Windows below that threshold are never uploaded, which is what keeps a quiet camera near zero requests, but a camera pointed at a busy street will keep sending.
Three things keep this opt-in: `transcribe` is not one of the default roles, live transcription is off by default, and the volume gate suppresses silence. Transcription of recorded `speech` events is unaffected - it remains a manual, one-request-per-event action.
:::
`device` and `model_size` have no effect on this path and no local model is ever downloaded.
`language` defaults to `auto`, which sends no language hint and lets the model detect it. Most audio models detect language well, so leave it on `auto` unless detection is picking the wrong one.
When set explicitly, it is sent as the transcription endpoint's native `language` parameter for OpenAI, Azure, and llama.cpp, and as part of the prompt for Gemini. This matters for dedicated ASR models such as Qwen3-ASR: they read the prompt as contextual biasing rather than as an instruction, so a language named in the prompt is ignored, while the endpoint parameter is honored.
The only field that is valid at the camera level is `enabled`.
#### Live transcription
@@ -362,8 +292,6 @@ Results can be error-prone due to a number of factors, including:
For speech sources close to the camera with minimal background noise, use the `small` model.
A [GenAI provider](#genai-provider) is generally the most accurate option for live transcription, at the cost of a network round trip per window. That round trip has to stay under about a second to keep up with the audio; if it does not, Frigate drops the oldest buffered audio rather than letting the backlog grow.
If you have CUDA hardware, you can experiment with the `large` `whisper` model on GPU. Performance is not quite as fast as the `sherpa-onnx` `small` model, but live transcription is far more accurate. Using the `large` model with CPU will likely be too slow for real-time transcription.
#### Transcription and translation of `speech` audio events
@@ -380,7 +308,7 @@ Only one `speech` event may be transcribed at a time. Frigate does not automatic
:::
With `model: whisper`, recorded `speech` events always use a `whisper` model, regardless of the `model_size` config setting. Without a supported Nvidia GPU, generating transcriptions for longer `speech` events may take a fair amount of time, so be patient. With a [GenAI provider](#genai-provider), the recorded clip is sent to the provider instead and no local model is used.
Recorded `speech` events will always use a `whisper` model, regardless of the `model_size` config setting. Without a supported Nvidia GPU, generating transcriptions for longer `speech` events may take a fair amount of time, so be patient.
#### FAQ
+6 -17
View File
@@ -43,7 +43,7 @@ genai:
The examples on this page all use `my_provider`, but the name is arbitrary and is only used to reference the provider elsewhere in the config (for example, `semantic_search.model`).
Each provider handles one or more **roles**: `chat`, `descriptions`, `embeddings`, and `transcribe`. A provider handles the first three by default; `transcribe` must always be listed explicitly, and is not available on Ollama, which has no audio input. Each role may be assigned to exactly one provider. Define a single provider if you want it to do everything, or split the roles across several providers using the `roles` option.
Each provider handles one or more **roles**: `chat`, `descriptions`, and `embeddings`. A provider handles all three by default, and each role may be assigned to exactly one provider. Define a single provider if you want it to do everything, or split the roles across several providers using the `roles` option.
If the provider you choose requires an API key, you may either directly paste it in your configuration, or store it in an environment variable prefixed with `FRIGATE_`.
@@ -63,11 +63,11 @@ Running Generative AI models on CPU is not recommended, as high inference times
You must use a vision-capable model with Frigate. The following models are recommended for local deployment of the `descriptions` and `chat` roles:
| Model | Review [frame mode](/configuration/genai/genai_review#frame-mode) | Notes |
| ------------------- | --------------------------------------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `qwen3-vl` | `frames` | Strong visual and situational understanding, enhanced ability to identify smaller objects and interactions with object. Follows a sequence of frames on its own. |
| `qwen3.6`/`qwen3.8` | `frames` | Strong situational understanding, but missing DeepStack from qwen3-vl leading to worse performance for identifying objects in people's hand and other small details. |
| `gemma4` | `annotated_frames` | Strong situational understanding, sometimes resorts to more vague terms like 'interacts' instead of assigning a specific action. Loses track of activity that repeats or reverses, so it benefits from annotated frames. |
| Model | Notes |
| ------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `qwen3-vl` | Strong visual and situational understanding, enhanced ability to identify smaller objects and interactions with object. |
| `qwen3.6`/`qwen3.8` | Strong situational understanding, but missing DeepStack from qwen3-vl leading to worse performance for identifying objects in people's hand and other small details. |
| `gemma4` | Strong situational understanding, sometimes resorts to more vague terms like 'interacts' instead of assigning a specific action. |
#### Embedding models
@@ -77,17 +77,6 @@ The `embeddings` role needs a different kind of model. Text queries are matched
| -------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `qwen3-vl-embedding` | Multimodal embeddings for [Semantic Search](/configuration/semantic_search#genai-provider). Must be served by llama.cpp started with `--embeddings` and `--mmproj`. |
#### Transcription models
The `transcribe` role needs a model that accepts audio input. A text-only or vision-only model cannot serve this role. The following are recommended for local deployment of the `transcribe` role:
| Model | Notes |
| ----------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
| `qwen3-asr` | Dedicated speech recognition model covering 30 languages, and the better choice for transcription quality. It only transcribes, so it cannot be shared with the `descriptions` or `chat` roles. |
| `gemma4` | General multimodal model that accepts audio as well as images, so one served model can cover `transcribe` alongside the other roles. Transcript quality is below `qwen3-asr`, particularly on noisy audio. |
Both must be served by llama.cpp started with the matching audio `--mmproj`. llama.cpp only reports audio support when an audio projector is loaded. Without it Frigate sees the model as text-only and the `transcribe` role is unavailable in the UI. Frigate transcribes through the server's `/v1/audio/transcriptions` route, which llama.cpp serves for any audio-capable model.
:::info
Each model is available in multiple parameter sizes (3b, 4b, 8b, etc.). Larger sizes are more capable of complex tasks and understanding of situations, but requires more memory and computational resources. It is recommended to try multiple models and experiment to see which performs best.
@@ -192,43 +192,6 @@ review:
</TabItem>
</ConfigTabs>
### Frame Mode
Review items are sent to the model as a sequence of still frames. Some models follow that sequence well on their own; others lose track of activity that repeats or reverses, and describe a single trip when the subject actually made several. The `frame_mode` option controls how those frames are presented.
- `frames` (default): the prompt followed by the frames, exactly as earlier versions of Frigate sent them.
- `annotated_frames`: each frame is preceded by its frame number and elapsed time, along with notes describing what the object tracker recorded at that moment, such as an object being first detected, starting to move, turning around, stopping, or no longer being detected.
The notes come from tracking data rather than from the images, so they describe activity the model may not have picked up on its own. In testing with a person carrying three waste bins to the curb one at a time, `gemma4` described a single trip on every attempt with `frames`, and consistently described multiple trips with `annotated_frames`. Models that already handle these sequences well, such as the `qwen3-vl` family, gain little and should stay on `frames`.
Annotated mode also caps the number of frames, since the notes already establish the order of events and extra near-duplicate frames tend to crowd out the middle of a clip. Longer review items are sampled more sparsely as a result, and typically use fewer tokens than `frames` mode for the same item.
:::note
Annotated mode needs tracking data for the review item. If none is available, Frigate falls back to sending plain frames for that item.
:::
<ConfigTabs>
<TabItem value="ui">
Navigate to <NavPath path="Settings > Global configuration > Review" />.
- Set **GenAI config > Frame mode** to the desired mode (e.g., `annotated_frames`)
</TabItem>
<TabItem value="yaml">
```yaml {4}
review:
genai:
enabled: true
frame_mode: annotated_frames
```
</TabItem>
</ConfigTabs>
### Response Style
Different models respond to the built-in prompt with very different writing styles: some produce natural narration while others sound short and mechanical. The `response_style` option selects a writing style preset that rewords the prompt's instructions for the user-facing fields (the title, short summary, and scene description). Presets replace those instructions rather than adding extra ones, so the model never receives competing style directions.
@@ -499,7 +499,7 @@ cameras:
## Synaptics
Hardware accelerated video de-/encoding is supported on Synaptics SL-series SoC.
Hardware accelerated video de-/encoding is supported on Synpatics SL-series SoC.
### Prerequisites
-31
View File
@@ -28,24 +28,6 @@ WebRTC may use an external STUN server for NAT traversal. MSE and HLS streaming
:::
### Selecting a streaming technology
Frigate [defaults to MSE](#why-does-frigate-prefer-mse-over-webrtc-for-live-view) for restreamed cameras by design. To use WebRTC, select it explicitly from a camera's single-camera Live view settings (the settings menu in the camera's Live view header on desktop, or the settings drawer on mobile). Three related controls work together:
- **Stream**: _what_ to play. This lists the [streams you've configured](#setting-streams-for-live-ui) (for example `Main Stream` and `Sub Stream`).
- **Force low-bandwidth mode**: a switch that always plays Frigate's built-in low-bandwidth feed (the stream assigned the `detect` role, using JSMpeg) instead of the selected stream. It works anywhere without go2rtc and is useful on slow or metered connections. While it is enabled, the stream and streaming technology selectors are disabled; your stream and technology choices are restored when you turn it off.
- **Streaming Technology**: _how_ to play the selected stream, listing **MSE** and **WebRTC**. It is only shown for a restreamed stream.
- The choices are saved **per device, per camera** in your browser's local storage.
- **WebRTC is only selectable when it can actually work for that stream.** When it can't, the option is shown disabled with the reason inline, and a more detailed reason (the failing codecs, or why the connectivity check failed) is logged to your browser's console. Common reasons:
- **Not configured**: no `candidates` or `ice_servers` are set under `go2rtc.webrtc` (see [WebRTC extra configuration](#webrtc-extra-configuration)).
- **Could not connect**: e.g. port `8555` isn't reachable, or a STUN/TURN server is misconfigured. Frigate runs a one-time WebRTC connectivity check when the Live view opens; the option may briefly show as "checking" while it runs.
- **Unsupported video codec**: the stream's video codec can't be played over WebRTC in your browser, most commonly H.265/HEVC in Firefox or Edge.
- **Unsupported audio codec**: WebRTC needs opus or G.711 audio, so a stream whose playback audio is only AAC (without an added opus/G.711 track) can't carry audio over WebRTC. See [Audio Support](#audio-support) for how to add one.
- **Unsupported browser**: the browser doesn't support WebRTC.
When WebRTC isn't available, Frigate automatically uses MSE (or falls back to JSMpeg), so live view keeps working regardless of the selection.
### Camera Settings Recommendations
If you are using go2rtc, you should adjust the following settings in your camera's firmware for the best experience with Live view:
@@ -175,17 +157,6 @@ WebRTC works by creating a TCP or UDP connection on port `8555`. However, it req
- stun:8555
```
- The web UI uses the STUN and TURN servers in `ice_servers` and falls back to Google's public STUN server when none are set:
```yaml title="config.yml"
go2rtc:
webrtc:
ice_servers:
- urls: [turn:turn.example.com:3478]
username: frigate
credential: password
```
- For access through Tailscale, the Frigate system's Tailscale IP must be added as a WebRTC candidate. Tailscale IPs all start with `100.`, and are reserved within the `100.64.0.0/10` CIDR block.
- Note that some browsers may not support H.265 (HEVC). You can check your browser's current version for H.265 compatibility [here](https://github.com/AlexxIT/go2rtc?tab=readme-ov-file#codecs-madness).
@@ -235,8 +206,6 @@ For devices that support two way talk, Frigate can be configured to use the feat
- Ensure you access Frigate via https (may require [opening port 8971](/frigate/installation/#ports)).
- For the Home Assistant Frigate card, [follow the docs](http://card.camera/#/usage/2-way-audio) for the correct source.
The two-way talk control in the single-camera Live view is only enabled when WebRTC is available; if WebRTC isn't configured or can't connect, the control is shown disabled.
To use the Reolink Doorbell with two way talk, you should use the [recommended Reolink configuration](/configuration/camera_specific#reolink-cameras)
As a starting point to check compatibility for your camera, view the list of cameras supported for two-way talk on the [go2rtc repository](https://github.com/AlexxIT/go2rtc?tab=readme-ov-file#two-way-audio). For cameras in the category `ONVIF Profile T`, you can use the [ONVIF Conformant Products Database](https://www.onvif.org/conformant-products/)'s FeatureList to check for the presence of `AudioOutput`. A camera that supports `ONVIF Profile T` _usually_ supports this, but due to inconsistent support, a camera that explicitly lists this feature may still not work. If no entry for your camera exists on the database, it is recommended not to buy it or to consult with the manufacturer's support on the feature availability.
-8
View File
@@ -313,16 +313,8 @@ To remove root from the container entirely, add Docker's `user:`:
```yaml
user: "1000:1000" # NOT compatible with PUID/PGID, see the run modes table
tmpfs:
- /tmp:size=256m
- /tmp/cache:size=1000000000
- /run:exec,nosuid,nodev,mode=0755,uid=1000,gid=1000,size=16m # uid must match user:
```
`/run` has to be owned by that uid as well. s6 writes its runtime state there before anything else starts, and with no root in the container a root-owned `/run` stops it during init with `cannot create /run/test of writability`. Keep `uid` and `gid` in the tmpfs options matching `user:`, and don't carry that pair back into the default mode, where a root-owned `/run` is what keeps the unprivileged services out of s6's runtime state.
This only bites once root is genuinely gone. s6's init helper is setuid, so `user:` on its own still lets init regain root and correct `/run` itself. The `no-new-privileges:true` above is what blocks that, which is also what makes the `/run` ownership mandatory. Dropping it would hide the problem by handing init root again.
Two things change, and the first one will break a working install if you skip it. The startup device grants can't run, because there is no root left to run them, so every device you pass stops working until you grant that uid access yourself with `group_add:` or a udev rule; see [Manual setup](#manual-setup). Expect this to surface as a driver error rather than a permission error, like `No VA display found` from VAAPI. And every service then runs as that one uid, so go2rtc no longer gets its own restricted user. `/config` and `/media/frigate` have to be owned by that uid already, since Frigate never adjusts ownership in this mode. Switching an existing install over also leaves `/config/go2rtc_homekit.yml` owned by the go2rtc user, which this mode can't write; `chown` it to your uid or HomeKit pairing changes stop persisting. Frigate warns and starts either way.
This mode can also take `cap_drop: [ALL]`, which the default mode cannot: starting as root needs `CAP_CHOWN` for the ownership sweep, `CAP_SETUID` and `CAP_SETGID` to drop to the runtime user, and `CAP_FOWNER` for the device grants.
+30 -6
View File
@@ -22,7 +22,7 @@ Frigate supports multiple different detectors that work on different types of ha
**Most Hardware**
- [Coral EdgeTPU](#edge-tpu-detector): The Google Coral EdgeTPU is available in USB, Mini PCIe, and m.2 formats allowing for a wide range of compatibility with devices.
- [Hailo](#hailo): The Hailo-8, Hailo-8L and Hailo-8R AI Acceleration modules are available in m.2 format with a HAT for RPi devices, offering a wide range of compatibility with devices.
- [Hailo](#hailo-8): The Hailo8 and Hailo8L AI Acceleration module is available in m.2 format with a HAT for RPi devices, offering a wide range of compatibility with devices.
- <CommunityBadge /> [MemryX](#memryx-mx3): The MX3 Acceleration module is available in m.2 format, offering broad compatibility across various platforms.
**AMD**
@@ -285,9 +285,9 @@ models:
---
## Hailo
## Hailo-8
This detector is available for use with the Hailo-8, Hailo-8L and Hailo-8R AI Acceleration Modules. The integration identifies which of them is attached and selects the matching default model if no custom model is specified.
This detector is available for use with both Hailo-8 and Hailo-8L AI Acceleration Modules. The integration automatically detects your hardware architecture via the Hailo CLI and selects the appropriate default model if no custom model is specified.
See the [installation docs](../frigate/installation.md#hailo-8) for information on configuring the Hailo hardware.
@@ -308,11 +308,11 @@ The HailoRT runtime is not part of the Frigate image. It is downloaded and insta
When configuring the Hailo detector, you have two options to specify the model: a local **path** or a **URL**.
If both are provided, the detector will first check for the model at the given local path. If the file is not found, it will download the model from the specified URL. The model file is cached under `/config/model_cache/hailo`.
<ModelConfigDropdown detectorTitle="Hailo" models={objectDetectorsModels.hailo.models} />
<ModelConfigDropdown detectorTitle="Hailo-8/Hailo-8L" models={objectDetectorsModels.hailo8l.models} />
For additional ready-to-use models, please visit: https://github.com/hailo-ai/hailo_model_zoo
Hailo supports all models in the Hailo Model Zoo that include HailoRT post-processing. You're welcome to choose any of these pre-configured models for your implementation.
Hailo8 supports all models in the Hailo Model Zoo that include HailoRT post-processing. You're welcome to choose any of these pre-configured models for your implementation.
> **Note:**
> The config.path parameter can accept either a local file path or a URL ending with .hef. When provided, the detector will first check if the path is a local file path. If the file exists locally, it will use it directly. If the file is not found locally or if a URL was provided, it will attempt to download the model from the specified URL.
@@ -362,7 +362,7 @@ Intel NPUs cannot be used under Home Assistant OS, which does not include the NP
:::warning
The Apple Silicon detector client is being reworked. Its extra options no longer have a place in the config, so only the endpoint carried in the device string is honored right now, and `request_timeout_ms` and `linger_ms` are ignored. Anything else is dropped when your config is migrated.
The network-based detectors (Deepstack and the Apple Silicon client) are being reworked. Their extra options no longer have a place in the config, so only the endpoint carried in the device string is honored right now: Deepstack ignores `api_key` and `api_timeout`, and the Apple Silicon client ignores `request_timeout_ms` and `linger_ms`. Anything else is dropped when your config is migrated.
:::
@@ -540,6 +540,30 @@ A TensorFlow Lite model is provided in the container at `/cpu_model.tflite` and
When using CPU detectors, you can add one CPU detector per camera. Adding more detectors than the number of cameras should not improve performance.
## Deepstack / CodeProject.AI Server Detector
:::warning
The network-based detectors (Deepstack and the Apple Silicon client) are being reworked. Their extra options no longer have a place in the config, so only the endpoint carried in the device string is honored right now: Deepstack ignores `api_key` and `api_timeout`, and the Apple Silicon client ignores `request_timeout_ms` and `linger_ms`. Anything else is dropped when your config is migrated.
:::
The Deepstack / CodeProject.AI Server detector for Frigate allows you to integrate Deepstack and CodeProject.AI object detection capabilities into Frigate. CodeProject.AI and DeepStack are open-source AI platforms that can be run on various devices such as the Raspberry Pi, Nvidia Jetson, and other compatible hardware. It is important to note that the integration is performed over the network, so the inference times may not be as fast as native Frigate detectors, but it still provides an efficient and reliable solution for object detection and tracking.
### Setup {#setup-deepstack}
To get started with CodeProject.AI, visit their [official website](https://www.codeproject.com/Articles/5322557/CodeProject-AI-Server-AI-the-easy-way) to follow the instructions to download and install the AI server on your preferred device. Detailed setup instructions for CodeProject.AI are outside the scope of the Frigate documentation.
To integrate CodeProject.AI into Frigate, configure the detector as follows:
### Configuration {#configuration-deepstack}
<ModelConfigDropdown detectorTitle="DeepStack" models={objectDetectorsModels.deepstack.models} />
Replace `<your_codeproject_ai_server_ip>` and `<port>` with the IP address and port of your CodeProject.AI server.
To verify that the integration is working correctly, start Frigate and observe the logs for any error messages related to CodeProject.AI. Additionally, you can check the Frigate web interface to see if the objects detected by CodeProject.AI are being displayed and tracked properly.
# Community Supported Detectors
## MemryX MX3
+1 -1
View File
@@ -450,7 +450,7 @@ For advanced use cases, the [custom export HTTP API](../integrations/api/export-
POST /export/custom/{camera_name}/start/{start_time}/end/{end_time}
```
The request body accepts `ffmpeg_input_args` and `ffmpeg_output_args` to control encoding, frame rate, filters, and other FFmpeg options. If neither is provided, Frigate defaults to time-lapse output settings (25x speed, 30 FPS) with audio removed (`-an`). When providing your own `ffmpeg_input_args`, include `-an` if you want audio stripped from the export.
The request body accepts `ffmpeg_input_args` and `ffmpeg_output_args` to control encoding, frame rate, filters, and other FFmpeg options. If neither is provided, Frigate defaults to time-lapse output settings (25x speed, 30 FPS).
The following example exports a time-lapse at 60x speed with 25 FPS:
+1 -3
View File
@@ -197,7 +197,7 @@ For cameras that support two-way talk, go2rtc will automatically establish an au
To prevent this, you must configure two separate stream instances:
1. One stream instance with `#backchannel=0` for Frigate's viewing, recording, and detection (prevents go2rtc from establishing the blocking backchannel)
2. A second stream instance with no `#` parameters at all for two-way talk functionality (can be used by Frigate's WebRTC viewer or other applications)
2. A second stream instance without `#backchannel=0` for two-way talk functionality (can be used by Frigate's WebRTC viewer or other applications)
Configuration example:
@@ -215,8 +215,6 @@ In this configuration:
- `front_door` stream is used by Frigate for viewing, recording, and detection. The `#backchannel=0` parameter prevents go2rtc from establishing the audio output backchannel, so it won't block two-way talk access.
- `front_door_twoway` stream is used for two-way talk functionality. This stream can be used by Frigate's WebRTC viewer when two-way talk is enabled, or by other applications (like Home Assistant Advanced Camera Card) that need access to the camera's audio output channel.
Any `#` parameter on a bare `rtsp://` source disables the backchannel unless the URL explicitly contains `#backchannel=1`. A two-way talk stream with something like `#video=h264` on it silently loses two-way audio, and Frigate will report that two-way talk is unavailable for that stream.
## Security: Restricted Stream Sources
For security reasons, the `echo:`, `expr:`, and `exec:` stream sources are disabled by default in go2rtc. These sources allow arbitrary command execution and can pose security risks if misconfigured.
+3 -12
View File
@@ -204,20 +204,11 @@ Light guidelines and advice:
npm run lint
```
- Ensure the backend [unit tests](#unit-tests) pass. Your PR cannot be merged unless tests pass.
```shell
python3 -u -m unittest
```
- Ensure the end-to-end tests pass. They run in Playwright against a production build with mocked API data, so they don't need a running Frigate instance. Add or update tests in `web/e2e/specs/` when you change UI behavior.
- Add to unit tests and ensure they pass. As much as possible, you should strive to _increase_ test coverage whenever making changes. This will help ensure features do not accidentally become broken in the future.
- If you run into error messages like "TypeError: Cannot read properties of undefined (reading 'context')" when running tests, this may be due to these issues (https://github.com/vitest-dev/vitest/issues/1910, https://github.com/vitest-dev/vitest/issues/1652) in vitest, but I haven't been able to resolve them.
```console
# First-time setup
npx playwright install chromium
# Build the app and run all tests
npm run e2e:build && npm run e2e
npm run test
```
- Test in different browsers. Firefox, Chrome, and Safari all have different quirks that make them unique targets to interact with.
+4 -5
View File
@@ -54,7 +54,7 @@ Frigate supports multiple different detectors that work on different types of ha
**Most Hardware**
- [Hailo](#hailo-8): The Hailo-8, Hailo-8L and Hailo-8R AI Acceleration modules are available in m.2 format with a HAT for RPi devices offering a wide range of compatibility with devices.
- [Hailo](#hailo-8): The Hailo8 and Hailo8L AI Acceleration module is available in m.2 format with a HAT for RPi devices offering a wide range of compatibility with devices.
- [Supports many model architectures](../../configuration/object_detectors#configuration-hailo)
- Runs best with tiny or small size models
@@ -111,13 +111,12 @@ Frigate supports multiple different detectors that work on different types of ha
### Hailo-8
Frigate supports the Hailo-8, Hailo-8L and Hailo-8R AI Acceleration Modules on compatible hardware platforms, including the Raspberry Pi 5 with the PCIe hat from the AI kit. The Hailo detector integration in Frigate identifies which of them is attached and selects the matching default model when a custom model isn’t provided.
Frigate supports both the Hailo-8 and Hailo-8L AI Acceleration Modules on compatible hardware platforms, including the Raspberry Pi 5 with the PCIe hat from the AI kit. The Hailo detector integration in Frigate automatically identifies your hardware type and selects the appropriate default model when a custom model isn’t provided.
**Default Model Configuration:**
- **Hailo-8L:** Default model is **YOLOv6n**, compiled for the Hailo-8L.
- **Hailo-8:** Default model is **YOLOv6n**, compiled for the Hailo-8.
- **Hailo-8R:** Default model is the **Hailo-8** build of **YOLOv6n**, since the Hailo Model Zoo publishes no Hailo-8R build.
- **Hailo-8L:** Default model is **YOLOv6n**.
- **Hailo-8:** Default model is **YOLOv6n**.
In real-world deployments, even with multiple cameras running concurrently, Frigate has demonstrated consistent performance. Testing on x86 platforms, with dual PCIe lanes, yields further improvements in FPS, throughput, and latency compared to the Raspberry Pi setup.
+2 -2
View File
@@ -122,7 +122,7 @@ Additionally, the USB Coral draws a considerable amount of power. If using any o
### Hailo-8
The Hailo-8, Hailo-8L and Hailo-8R AI accelerators are available in both M.2 and HAT form factors for the Raspberry Pi. The M.2 version typically connects to a carrier board for PCIe, which then interfaces with the Raspberry Pi 5 as part of the AI Kit. The HAT version can be mounted directly onto compatible Raspberry Pi models. Both form factors have been successfully tested on x86 platforms as well, making them versatile options for various computing environments.
The Hailo-8 and Hailo-8L AI accelerators are available in both M.2 and HAT form factors for the Raspberry Pi. The M.2 version typically connects to a carrier board for PCIe, which then interfaces with the Raspberry Pi 5 as part of the AI Kit. The HAT version can be mounted directly onto compatible Raspberry Pi models. Both form factors have been successfully tested on x86 platforms as well, making them versatile options for various computing environments.
The HailoRT runtime is not part of the Frigate image; Frigate downloads and installs it at first start once a Hailo detector is configured. Containers without internet access can provide the files themselves, see [Detector runtimes](/frigate/network_requirements#detector-runtimes).
@@ -300,7 +300,7 @@ If you are using `docker run`, add this option to your command `--device /dev/ha
#### Configuration
Finally, configure [hardware object detection](/configuration/object_detectors#hailo) to complete the setup.
Finally, configure [hardware object detection](/configuration/object_detectors#hailo-8) to complete the setup.
### MemryX MX3
+12 -8
View File
@@ -47,7 +47,7 @@ If you are using one of the following hardware detectors and have not provided y
| Detector | Model Downloaded | Source |
| ------------------------------------------------------------------ | -------------------- | ------------------------ |
| [Rockchip RKNN](/configuration/object_detectors#rockchip-platform) | RKNN detection model | GitHub |
| [Hailo 8 / 8L / 8R](/configuration/object_detectors#hailo) | YOLOv6n (.hef) | Hailo Model Zoo (AWS S3) |
| [Hailo 8 / 8L](/configuration/object_detectors#hailo-8) | YOLOv6n (.hef) | Hailo Model Zoo (AWS S3) |
| [AXERA AXEngine](/configuration/object_detectors) | Detection model | HuggingFace |
:::note
@@ -60,16 +60,16 @@ The default CPU, EdgeTPU, and OpenVINO object detection models are bundled into
The SDKs for a few hardware detectors are not shipped in the Frigate image. They are downloaded the first time that detector is configured, verified against checksums pinned in the Frigate release, and installed into the Frigate user's home directory (`/config/.local` by default). Once installed they are not downloaded again until a Frigate release pins a new version.
| Detector | Version | Files | Source |
| ---------------------------------------------------------- | ------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ---------------------------------------------------------------------------------------- |
| [Hailo 8 / 8L / 8R](/configuration/object_detectors#hailo) | 4.21.0 | `hailort-debian12-amd64.tar.gz` and `hailort-4.21.0-cp311-cp311-linux_x86_64.whl` on x86, `hailort-debian12-arm64.tar.gz` and `hailort-4.21.0-cp311-cp311-linux_aarch64.whl` on arm64 | [GitHub release](https://github.com/frigate-nvr/hailort/releases/tag/v4.21.0) |
| [MemryX MX3](/configuration/object_detectors#memryx-mx3) | 2.1.0 | `mx_accl_frigate-2.1.0.zip` (the release source archive, renamed) | [GitHub archive](https://github.com/memryx/mx_accl_frigate/archive/refs/tags/v2.1.0.zip) |
| [AXERA AXEngine](/configuration/object_detectors#axera) | 0.1.3 | `axengine-0.1.3-py3-none-any.whl` | [GitHub release](https://github.com/AXERA-TECH/pyaxengine/releases/tag/0.1.3-frigate) |
| Detector | Version | Files | Source |
| -------------------------------------------------------------- | ------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------ |
| [Hailo 8 / 8L](/configuration/object_detectors#hailo-8) | 4.21.0 | `hailort-debian12-amd64.tar.gz` and `hailort-4.21.0-cp311-cp311-linux_x86_64.whl` on x86, `hailort-debian12-arm64.tar.gz` and `hailort-4.21.0-cp311-cp311-linux_aarch64.whl` on arm64 | [GitHub release](https://github.com/frigate-nvr/hailort/releases/tag/v4.21.0) |
| [MemryX MX3](/configuration/object_detectors#memryx-mx3) | 2.1.0 | `mx_accl_frigate-2.1.0.zip` (the release source archive, renamed) | [GitHub archive](https://github.com/memryx/mx_accl_frigate/archive/refs/tags/v2.1.0.zip) |
| [AXERA AXEngine](/configuration/object_detectors#axera) | 0.1.3 | `axengine-0.1.3-py3-none-any.whl` | [GitHub release](https://github.com/AXERA-TECH/pyaxengine/releases/tag/0.1.3-frigate) |
If the container cannot reach GitHub, provide the files yourself:
1. Download the files for your architecture on a machine with internet access.
2. Place them, with exactly the file names listed above, in `/config/model_cache/runtimes/<detector>/`, where `<detector>` is the detector named in your config's `devices` (`hailo`, `memryx`, or `axengine`).
2. Place them, with exactly the file names listed above, in `/config/model_cache/runtimes/<detector>/`, where `<detector>` is the detector `type` from your config (`hailo8l`, `memryx`, or `axengine`).
3. Start Frigate. Files whose checksum matches are installed without any download; a file with the wrong checksum is discarded and downloaded again, so a failed startup log names the file to replace.
The `GITHUB_ENDPOINT` mirror variable below applies to these downloads as well.
@@ -142,12 +142,16 @@ When [notifications](/configuration/notifications) are enabled and users have re
If an [MQTT broker](/integrations/mqtt) is configured, Frigate maintains a connection to the broker's host and port. This is typically a local network connection, but will require internet if you use a cloud-hosted MQTT broker.
### DeepStack / CodeProject.AI
When using the [DeepStack detector plugin](/configuration/object_detectors), Frigate sends images to the configured API endpoint for inference. This is typically local but depends on where the service is hosted.
## WebRTC (STUN)
For [WebRTC live streaming](/configuration/live), Frigate uses STUN for NAT traversal:
- **go2rtc** defaults to a local STUN listener (`stun:8555`), no internet required.
- **The web UI** uses the servers in `go2rtc.webrtc.ice_servers` for its WebRTC player and for the WebRTC connectivity check it runs when the Live view loads. If none are set, it uses Google's public STUN server (`stun:stun.l.google.com:19302`), which requires internet access from the browser. Set `ice_servers` to a STUN or TURN server on your network to avoid this.
- **The web UI's WebRTC player** includes a fallback to Google's public STUN server (`stun:stun.l.google.com:19302`), which requires internet.
## Home Assistant Supervisor
+2 -2
View File
@@ -41,7 +41,7 @@ Rockchip models are automatically converted as of 0.17. For 0.16, YOLOv9 onnx mo
## Supported detector types
Currently, Frigate+ models support CPU (`cpu`), Google Coral (`edgetpu`), OpenVino (`openvino`), ONNX (`onnx`), Hailo (`hailo`), and Rockchip (`rknn`) detectors.
Currently, Frigate+ models support CPU (`cpu`), Google Coral (`edgetpu`), OpenVino (`openvino`), ONNX (`onnx`), Hailo (`hailo8l`), and Rockchip (`rknn`) detectors.
| Hardware | Recommended Detector Type | Recommended Model Type |
| -------------------------------------------------------------------------------- | ------------------------- | ---------------------- |
@@ -50,7 +50,7 @@ Currently, Frigate+ models support CPU (`cpu`), Google Coral (`edgetpu`), OpenVi
| [Intel](/configuration/object_detectors.md#openvino-detector) | `openvino` | `yolov9` |
| [NVidia GPU](/configuration/object_detectors#onnx) | `onnx` | `yolov9` |
| [AMD ROCm GPU](/configuration/object_detectors#amdrocm-gpu-detector) | `onnx` | `yolov9` |
| [Hailo8/Hailo8L/Hailo8R](/configuration/object_detectors#hailo) | `hailo` | `yolov9` |
| [Hailo8/Hailo8L/Hailo8R](/configuration/object_detectors#hailo-8) | `hailo8l` | `yolov9` |
| [Rockchip NPU](/configuration/object_detectors#rockchip-platform) | `rknn` | `yolov9` |
## Improving your model
-14
View File
@@ -39,20 +39,6 @@ To do this efficiently the following setup is required:
When this is done correctly, the GPU will do the decoding and scaling which will result in a small increase in CPU usage but with better results.
### How can I rotate my camera's video feed?
Rotation is best done in the camera's firmware settings (usually called rotate, flip, or corridor mode) so the video arrives already rotated and no extra processing is needed. Check there first.
If your camera does not support rotation, go2rtc's ffmpeg module can rotate the stream with the `#rotate` parameter (`90`, `180`, `270`, or `-90`), but this is not recommended: rotation requires transcoding (re-encoding) the video, which significantly increases CPU usage, especially for high resolution streams.
```yaml
go2rtc:
streams:
my_camera: "ffmpeg:rtsp://user:password@192.168.1.10:554/stream#video=h264#hardware#rotate=90"
```
Point the camera's inputs at the restream as described in the [restream docs](/configuration/restream.md), and swap `detect -> width` and `detect -> height` to match the rotated resolution.
### My mjpeg stream or snapshots look green and crazy
This almost always means that the width/height defined for your camera are not correct. Double check the resolution with VLC or another player. Also make sure you don't have the width and height values backwards.
+2 -4
View File
@@ -78,9 +78,7 @@ go2rtc:
:::warning
The transcoding modifiers (`#video=`, `#audio=`, `#hardware`, …) **only take effect on a source that is prefixed with `ffmpeg:`**. Adding them to a bare `rtsp://…#audio=opus` source does nothing: go2rtc ignores them. Likewise, when a source references another stream by name (e.g. `ffmpeg:back#audio=aac`), the name must match the stream key **exactly** (it is case sensitive), or the transcode is silently never produced. This is the single most common configuration mistake. In the Frigate UI, the **Use compatibility mode (ffmpeg)** toggle adds the `ffmpeg:` prefix for you.
A bare `rtsp://` source reads a different set of modifiers: `#backchannel=`, `#media=`, `#timeout=`, and `#transport=`. These do nothing on an `ffmpeg:` source. Adding **any** modifier to a bare `rtsp://` source also disables the camera's backchannel unless the URL explicitly contains `#backchannel=1`, so a stream dedicated to two-way talk should carry no modifiers at all.
The `#`-modifiers (`#video=`, `#audio=`, `#hardware`, `#backchannel=0`, …) **only take effect on a source that is prefixed with `ffmpeg:`**. Adding them to a bare `rtsp://…#audio=opus` source does nothing: go2rtc ignores them. Likewise, when a source references another stream by name (e.g. `ffmpeg:back#audio=aac`), the name must match the stream key **exactly** (it is case sensitive), or the transcode is silently never produced. This is the single most common configuration mistake. In the Frigate UI, the **Use compatibility mode (ffmpeg)** toggle adds the `ffmpeg:` prefix for you.
:::
@@ -155,7 +153,7 @@ WebRTC is only attempted when MSE fails or when using a camera's two-way talk fe
- **Codec mismatch**: WebRTC cannot carry H.265 or AAC. The stream backing the WebRTC view must provide Opus (or PCMA/PCMU) audio and H.264 video. Add an `ffmpeg:back#audio=opus` source as shown above.
- **Port `8555` not reachable, or no candidates set**: WebRTC needs port `8555` (both TCP and UDP) open and a reachable candidate advertised. On Docker installs running on a custom/overlay network, go2rtc may advertise unreachable container IPs as ICE candidates; setting `webrtc.filters.candidates: []` and supplying only your host's LAN IP resolves this. See [WebRTC extra configuration](/configuration/live#webrtc-extra-configuration).
- **Two-way talk** additionally requires a secure context (HTTPS or the authenticated port `8971`, because browsers block microphone access on plain HTTP). The camera's RTSP backchannel must also be handled correctly: go2rtc seizes the backchannel by default, which blocks two-way audio for other consumers and can inject static. Disable it on the primary stream with `#backchannel=0` and use a separate dedicated stream for talk, carrying no `#` modifiers of any kind, as documented in [preventing go2rtc from blocking two-way audio](/configuration/restream#two-way-talk-restream).
- **Two-way talk** additionally requires a secure context (HTTPS or the authenticated port `8971`, because browsers block microphone access on plain HTTP). The camera's RTSP backchannel must also be handled correctly: go2rtc seizes the backchannel by default, which blocks two-way audio for other consumers and can inject static. Disable it on the primary stream with `#backchannel=0` and use a separate dedicated stream for talk, as documented in [preventing go2rtc from blocking two-way audio](/configuration/restream#two-way-talk-restream).
## High CPU usage
+3905 -2073
View File
File diff suppressed because it is too large Load Diff
+10 -10
View File
@@ -18,17 +18,17 @@
"write-heading-ids": "docusaurus write-heading-ids"
},
"dependencies": {
"@docusaurus/core": "^3.10.2",
"@docusaurus/plugin-content-docs": "^3.10.2",
"@docusaurus/preset-classic": "^3.10.2",
"@docusaurus/theme-mermaid": "^3.10.2",
"@docusaurus/core": "^3.7.0",
"@docusaurus/plugin-content-docs": "^3.7.0",
"@docusaurus/preset-classic": "^3.7.0",
"@docusaurus/theme-mermaid": "^3.7.0",
"@inkeep/docusaurus": "^2.0.16",
"@mdx-js/react": "^3.1.0",
"@types/js-yaml": "^4.0.9",
"clsx": "^2.1.1",
"docusaurus-plugin-openapi-docs": "^5.2.0",
"docusaurus-theme-openapi-docs": "^5.2.0",
"js-yaml": "^4.3.2",
"docusaurus-plugin-openapi-docs": "^4.5.1",
"docusaurus-theme-openapi-docs": "^4.5.1",
"js-yaml": "^4.1.1",
"marked": "^16.4.2",
"prism-react-renderer": "^2.4.1",
"raw-loader": "^4.0.2",
@@ -48,11 +48,11 @@
]
},
"devDependencies": {
"@docusaurus/module-type-aliases": "^3.10.2",
"@docusaurus/types": "^3.10.2",
"@docusaurus/module-type-aliases": "^3.7.0",
"@docusaurus/types": "^3.7.0",
"@types/react": "^18.3.27"
},
"engines": {
"node": ">=20.19"
"node": ">=18.0"
}
}
-4
View File
@@ -99,10 +99,6 @@ def main() -> None:
print("*** End Config Validation Errors ***")
print("*************************************************************")
# force a non-zero exit code for config failures
if args.validate_config:
sys.exit(1)
# attempt to start Frigate in recovery mode
try:
config = FrigateConfig.load(install=True, safe_load=True)
+9 -2
View File
@@ -2,6 +2,7 @@
import asyncio
import copy
import json
import logging
import os
import platform
@@ -10,6 +11,7 @@ import urllib
from datetime import datetime, timedelta
from functools import reduce
from io import StringIO
from pathlib import Path as FilePath
from typing import Any
import aiofiles
@@ -54,7 +56,6 @@ from frigate.jobs.media_sync import (
start_media_sync_job,
)
from frigate.models import Event, Timeline
from frigate.plus import load_plus_model_info
from frigate.stats.prometheus import get_metrics, update_metrics
from frigate.types import JobStatusTypesEnum
from frigate.util.builtin import (
@@ -400,7 +401,13 @@ def config(request: Request):
model_dict["plus"] = None
if model.path:
model_dict["plus"] = load_plus_model_info(os.path.basename(model.path))
model_json_path = FilePath(model.path).with_suffix(".json")
try:
with open(model_json_path) as f:
model_dict["plus"] = json.load(f)
except (FileNotFoundError, json.JSONDecodeError):
pass
return JSONResponse(content=config)
-7
View File
@@ -786,13 +786,6 @@ def auth(request: Request):
user = token.claims.get("sub")
role = token.claims.get("role")
# the token keeps the role it was issued with, so a role removed from
# the config since then must send the user back through login
if role not in auth_config.roles:
logger.debug("jwt role %s is not in the config", role)
return fail_response
current_time = int(time.time())
# if the jwt is expired
+5 -16
View File
@@ -71,7 +71,6 @@ from frigate.jobs.export import (
from frigate.models import Export, ExportCase, Previews, Recordings
from frigate.record.export import (
DEFAULT_TIME_LAPSE_FFMPEG_ARGS,
DEFAULT_TIME_LAPSE_FFMPEG_INPUT_ARGS,
ChaptersEnum,
ExportStreamEnum,
PlaybackSourceEnum,
@@ -1038,10 +1037,6 @@ def export_recording_custom(
if camera_validation_error is not None:
return camera_validation_error
# Validate user-provided ffmpeg args to prevent injection and add to cases.
# Admin users are trusted and skip validation.
is_admin = request.headers.get("remote-role", "") == "admin"
playback_source = body.source
friendly_name = body.name
existing_image, image_validation_error = _sanitize_existing_image(body.image_path)
@@ -1052,16 +1047,6 @@ def export_recording_custom(
cpu_fallback = body.cpu_fallback
export_case_id = body.export_case_id
if export_case_id is not None and not is_admin:
return JSONResponse(
content={
"success": False,
"message": "Only admins can attach exports to an existing case.",
},
status_code=403,
)
case_validation_error = _validate_export_case(export_case_id)
if case_validation_error is not None:
return case_validation_error
@@ -1078,6 +1063,10 @@ def export_recording_custom(
status_code=400,
)
# Validate user-provided ffmpeg args to prevent injection.
# Admin users are trusted and skip validation.
is_admin = request.headers.get("remote-role", "") == "admin"
if not is_admin:
for args_label, args_value in [
("input", ffmpeg_input_args),
@@ -1098,7 +1087,7 @@ def export_recording_custom(
# Set default values if not provided (timelapse defaults)
if ffmpeg_input_args is None:
ffmpeg_input_args = DEFAULT_TIME_LAPSE_FFMPEG_INPUT_ARGS
ffmpeg_input_args = ""
if ffmpeg_output_args is None:
ffmpeg_output_args = DEFAULT_TIME_LAPSE_FFMPEG_ARGS
+62 -83
View File
@@ -189,7 +189,7 @@ async def camera_ptz_info(request: Request, camera_name: str):
future = asyncio.run_coroutine_threadsafe(
request.app.onvif.get_camera_info(camera_name), request.app.onvif.loop
)
result = await asyncio.wrap_future(future)
result = future.result()
return JSONResponse(content=result)
else:
return JSONResponse(
@@ -258,15 +258,15 @@ async def latest_frame(
frame = request.app.camera_error_image
height = int(params.height or str(frame.shape[0]))
width = int(height * frame.shape[1] / frame.shape[0])
if frame is None:
return JSONResponse(
content={"success": False, "message": "Unable to get valid frame"},
status_code=500,
)
height = int(params.height or str(frame.shape[0]))
width = int(height * frame.shape[1] / frame.shape[0])
if height < 1 or width < 1:
return JSONResponse(
content="Invalid height / width requested :: {} / {}".format(
@@ -886,8 +886,9 @@ async def vod_event(
# If the recordings are not found and the event started more than 5 minutes ago, set has_clip to false
if (
event.start_time < datetime.now().timestamp() - 300
and isinstance(vod_response, JSONResponse)
and vod_response.status_code == 404
and type(vod_response) is tuple
and len(vod_response) == 2
and vod_response[1] == 404
):
Event.update(has_clip=False).where(Event.id == event_id).execute()
@@ -955,80 +956,64 @@ async def event_snapshot(
event_complete = False
jpg_bytes = None
frame_time = 0
try:
event = Event.get(Event.id == event_id, Event.end_time != None)
await require_camera_access(event.camera, request=request)
except DoesNotExist:
event = None
if event is not None:
event_complete = True
await require_camera_access(event.camera, request=request)
if not event.has_snapshot:
return JSONResponse(
content={"success": False, "message": "Snapshot not available"},
status_code=404,
)
snapshot_settings = _resolve_snapshot_settings(
request.app.frigate_config.cameras[event.camera].snapshots, params
)
jpg_bytes, frame_time = get_event_snapshot_bytes(
event,
ext="jpg",
timestamp=snapshot_settings["timestamp"],
bounding_box=snapshot_settings["bounding_box"],
crop=snapshot_settings["crop"],
height=snapshot_settings["height"],
quality=snapshot_settings["quality"],
timestamp_style=request.app.frigate_config.cameras[
event.camera
].timestamp_style,
colormap=request.app.frigate_config.model_for_camera(event.camera).colormap,
)
except DoesNotExist:
# see if the object is currently being tracked
try:
snapshot_settings = _resolve_snapshot_settings(
request.app.frigate_config.cameras[event.camera].snapshots, params
)
jpg_bytes, frame_time = get_event_snapshot_bytes(
event,
ext="jpg",
timestamp=snapshot_settings["timestamp"],
bounding_box=snapshot_settings["bounding_box"],
crop=snapshot_settings["crop"],
height=snapshot_settings["height"],
quality=snapshot_settings["quality"],
timestamp_style=request.app.frigate_config.cameras[
event.camera
].timestamp_style,
colormap=request.app.frigate_config.model_for_camera(
event.camera
).colormap,
camera_states: list[CameraState] = (
request.app.detected_frames_processor.get_camera_states()
)
for camera_state in camera_states:
if event_id in camera_state.tracked_objects:
tracked_obj = camera_state.tracked_objects.get(event_id)
if tracked_obj is not None:
snapshot_settings = _resolve_snapshot_settings(
camera_state.camera_config.snapshots, params
)
jpg_bytes, frame_time = tracked_obj.get_img_bytes(
ext="jpg",
timestamp=snapshot_settings["timestamp"],
bounding_box=snapshot_settings["bounding_box"],
crop=snapshot_settings["crop"],
height=snapshot_settings["height"],
quality=snapshot_settings["quality"],
)
await require_camera_access(camera_state.name, request=request)
except Exception:
return JSONResponse(
content={"success": False, "message": "Unknown error occurred"},
content={"success": False, "message": "Ongoing event not found"},
status_code=404,
)
else:
# see if the object is currently being tracked
camera_states: list[CameraState] = (
request.app.detected_frames_processor.get_camera_states()
except Exception:
return JSONResponse(
content={"success": False, "message": "Unknown error occurred"},
status_code=404,
)
for camera_state in camera_states:
tracked_obj = camera_state.tracked_objects.get(event_id)
if tracked_obj is None:
continue
await require_camera_access(camera_state.name, request=request)
try:
snapshot_settings = _resolve_snapshot_settings(
camera_state.camera_config.snapshots, params
)
jpg_bytes, frame_time = tracked_obj.get_img_bytes(
ext="jpg",
timestamp=snapshot_settings["timestamp"],
bounding_box=snapshot_settings["bounding_box"],
crop=snapshot_settings["crop"],
height=snapshot_settings["height"],
quality=snapshot_settings["quality"],
)
except Exception:
return JSONResponse(
content={"success": False, "message": "Ongoing event not found"},
status_code=404,
)
break
if jpg_bytes is None:
return JSONResponse(
content={"success": False, "message": "Live frame not available"},
@@ -1077,25 +1062,19 @@ async def event_thumbnail(
if not thumbnail_bytes:
# see if the object is currently being tracked
camera_states = request.app.detected_frames_processor.get_camera_states()
for camera_state in camera_states:
tracked_obj = camera_state.tracked_objects.get(event_id)
if tracked_obj is None:
continue
await require_camera_access(camera_state.name, request=request)
try:
thumbnail_bytes = tracked_obj.get_thumbnail(extension.value)
except Exception:
return JSONResponse(
content={"success": False, "message": "Event not found"},
status_code=404,
)
break
try:
camera_states = request.app.detected_frames_processor.get_camera_states()
for camera_state in camera_states:
if event_id in camera_state.tracked_objects:
tracked_obj = camera_state.tracked_objects.get(event_id)
if tracked_obj is not None:
await require_camera_access(camera_state.name, request=request)
thumbnail_bytes = tracked_obj.get_thumbnail(extension.value)
except Exception:
return JSONResponse(
content={"success": False, "message": "Event not found"},
status_code=404,
)
if not thumbnail_bytes:
return JSONResponse(
+5 -2
View File
@@ -412,8 +412,11 @@ async def no_recordings(
if not camera_list:
return JSONResponse(content=[])
before = params.before or datetime.now().timestamp()
after = params.after or (datetime.now() - timedelta(hours=1)).timestamp()
before = params.before or datetime.datetime.now().timestamp()
after = (
params.after
or (datetime.datetime.now() - datetime.timedelta(hours=1)).timestamp()
)
scale = params.scale
recordings: list[tuple[float, float]] = []
+2 -12
View File
@@ -483,7 +483,7 @@ class FrigateApp:
def start_audio_processor(self) -> None:
self.audio_process = AudioProcessor(
self.config, self.camera_metrics, self.embeddings_metrics, self.stop_event
self.config, self.camera_metrics, self.stop_event
)
self.audio_process.start()
self.processes["audio_detector"] = self.audio_process.pid or 0
@@ -556,20 +556,10 @@ class FrigateApp:
"output",
lambda: OutputProcess(self.config, self.stop_event),
),
(
"audio_process",
"audio_detector",
lambda: AudioProcessor(
self.config,
self.camera_metrics,
self.embeddings_metrics,
self.stop_event,
),
),
]
for attr, key, factory in specs:
if getattr(self, attr, None) is None:
if not hasattr(self, attr):
continue
def on_restart(
+3 -28
View File
@@ -1,7 +1,7 @@
from enum import Enum
from typing import Any, Self
from typing import Any
from pydantic import Field, model_validator
from pydantic import Field
from ..base import FrigateBaseModel
from ..env import EnvString
@@ -21,17 +21,6 @@ class GenAIRoleEnum(str, Enum):
chat = "chat"
descriptions = "descriptions"
embeddings = "embeddings"
transcribe = "transcribe"
# Providers that can accept audio input for the transcribe role. Ollama has no
# audio input support, so claiming the role there would fail at request time.
TRANSCRIBE_CAPABLE_PROVIDERS = {
GenAIProviderEnum.openai,
GenAIProviderEnum.azure_openai,
GenAIProviderEnum.gemini,
GenAIProviderEnum.llamacpp,
}
class GenAIConfig(FrigateBaseModel):
@@ -63,7 +52,7 @@ class GenAIConfig(FrigateBaseModel):
GenAIRoleEnum.chat,
],
title="Roles",
description="GenAI roles (chat, descriptions, embeddings, transcribe); one provider per role. Only chat, descriptions, and embeddings are granted by default; transcribe must be listed explicitly.",
description="GenAI roles (chat, descriptions, embeddings); one provider per role.",
)
provider_options: dict[str, Any] = Field(
default={},
@@ -77,17 +66,3 @@ class GenAIConfig(FrigateBaseModel):
description="Runtime options passed to the provider for each inference call.",
json_schema_extra={"additionalProperties": {}},
)
@model_validator(mode="after")
def validate_transcribe_provider(self) -> Self:
"""Reject the transcribe role on providers that cannot accept audio input."""
if (
GenAIRoleEnum.transcribe in self.roles
and self.provider not in TRANSCRIBE_CAPABLE_PROVIDERS
):
raise ValueError(
f"GenAI provider '{self.provider.value}' does not support audio input "
"and cannot be given the 'transcribe' role."
)
return self
-13
View File
@@ -9,7 +9,6 @@ __all__ = [
"DetectionsConfig",
"AlertsConfig",
"ImageSourceEnum",
"ReviewFrameModeEnum",
"ReviewResponseStyleEnum",
]
@@ -21,13 +20,6 @@ class ImageSourceEnum(str, Enum):
recordings = "recordings"
class ReviewFrameModeEnum(str, Enum):
"""How review frames are presented to the GenAI provider."""
frames = "frames"
annotated_frames = "annotated_frames"
class ReviewResponseStyleEnum(str, Enum):
"""Writing style presets for GenAI review descriptions."""
@@ -161,11 +153,6 @@ class GenAIReviewConfig(FrigateBaseModel):
description="Preferred language to request from the GenAI provider for generated responses.",
default=None,
)
frame_mode: ReviewFrameModeEnum = Field(
default=ReviewFrameModeEnum.frames,
title="Frame mode",
description="How frames are presented to the model. 'frames' sends the prompt followed by the frames, which suits models that track a sequence well on their own. 'annotated_frames' labels each frame and interleaves notes derived from object tracking, which helps models that lose track of activity that repeats or reverses.",
)
response_style: ReviewResponseStyleEnum = Field(
default=ReviewResponseStyleEnum.default,
title="Response style",
+2 -32
View File
@@ -5,7 +5,6 @@ from pydantic import ConfigDict, Field, field_validator
from .base import FrigateBaseModel
__all__ = [
"AudioTranscriptionModelEnum",
"CameraFaceRecognitionConfig",
"CameraLicensePlateRecognitionConfig",
"CameraAudioTranscriptionConfig",
@@ -21,10 +20,6 @@ class SemanticSearchModelEnum(str, Enum):
jinav2 = "jinav2"
class AudioTranscriptionModelEnum(str, Enum):
whisper = "whisper"
class EnrichmentsDeviceEnum(str, Enum):
GPU = "GPU"
CPU = "CPU"
@@ -58,35 +53,10 @@ class AudioTranscriptionConfig(FrigateBaseModel):
description="Enable or disable automatic audio transcription for all cameras; can be overridden per-camera.",
)
language: str = Field(
default="auto",
default="en",
title="Transcription language",
description="Language code used for transcription/translation (for example 'en' for English), or 'auto' to let the model detect it. See https://whisper-api.com/docs/languages/ for supported language codes.",
description="Language code used for transcription/translation (for example 'en' for English). See https://whisper-api.com/docs/languages/ for supported language codes.",
)
model: AudioTranscriptionModelEnum | str | None = Field(
default=AudioTranscriptionModelEnum.whisper,
title="Audio transcription model or GenAI provider name",
description="The transcription backend: 'whisper' for Frigate's built-in local models, or the name of a GenAI provider with the transcribe role.",
)
@field_validator("model", mode="before")
@classmethod
def coerce_model_enum(cls, v):
# An absent value ("model:" with nothing after it, or an explicit null)
# means unspecified, so fall back to the built-in backend. Left as None
# it would pass the GenAI-provider validation, which only inspects
# strings, and then be treated as a provider name that resolves to no
# client, turning transcription into a silent no-op.
if v is None or (isinstance(v, str) and not v.strip()):
return AudioTranscriptionModelEnum.whisper
if isinstance(v, str):
try:
return AudioTranscriptionModelEnum(v)
except ValueError:
return v
return v
device: EnrichmentsDeviceEnum = Field(
default=EnrichmentsDeviceEnum.CPU,
title="Transcription device",
+1 -31
View File
@@ -56,7 +56,6 @@ from .camera.timestamp import TimestampStyleConfig
from .camera_group import CameraGroupConfig
from .classification import (
AudioTranscriptionConfig,
AudioTranscriptionModelEnum,
ClassificationConfig,
FaceRecognitionConfig,
LicensePlateRecognitionConfig,
@@ -885,7 +884,7 @@ class FrigateConfig(FrigateBaseModel):
# set notifications state
self.notifications.enabled_in_config = self.notifications.enabled
# validate genai: each role (chat, descriptions, embeddings, transcribe) at most once
# validate genai: each role (chat, descriptions, embeddings) at most once
role_to_name: dict[GenAIRoleEnum, str] = {}
for name, genai_cfg in self.genai.items():
for role in genai_cfg.roles:
@@ -1246,35 +1245,6 @@ class FrigateConfig(FrigateBaseModel):
for model in self.models:
model.create_colormap(colored_labels)
# validate audio_transcription.model when it is a GenAI provider name.
# this runs here rather than beside the semantic_search check because the
# global->camera merge above is what resolves camera-level enablement.
transcription_active = self.audio_transcription.enabled or any(
camera.audio_transcription.enabled for camera in self.cameras.values()
)
if (
transcription_active
and isinstance(self.audio_transcription.model, str)
and not isinstance(
self.audio_transcription.model, AudioTranscriptionModelEnum
)
):
if self.audio_transcription.model not in self.genai:
raise ValueError(
f"audio_transcription.model '{self.audio_transcription.model}' is not a "
"valid GenAI config key. Must match a key in genai config."
)
if (
GenAIRoleEnum.transcribe
not in self.genai[self.audio_transcription.model].roles
):
raise ValueError(
f"GenAI provider '{self.audio_transcription.model}' must have "
"'transcribe' in its roles for audio transcription."
)
# Check audio transcription and audio detection requirements
if self.audio_transcription.enabled:
# If audio transcription is enabled globally, at least one camera must have audio detection enabled
@@ -10,7 +10,6 @@ from peewee import DoesNotExist
from frigate.comms.inter_process import InterProcessRequestor
from frigate.config import FrigateConfig
from frigate.config.classification import AudioTranscriptionModelEnum
from frigate.const import (
CACHE_DIR,
MODEL_CACHE_DIR,
@@ -19,13 +18,8 @@ from frigate.const import (
)
from frigate.data_processing.types import PostProcessDataEnum
from frigate.embeddings.embeddings import Embeddings
from frigate.genai.manager import GenAIClientManager
from frigate.types import TrackedObjectUpdateTypesEnum
from frigate.util.audio import (
clean_transcript,
get_audio_from_recording,
resolve_language,
)
from frigate.util.audio import get_audio_from_recording
from ..types import DataProcessorMetrics
from .api import PostProcessorApi
@@ -40,25 +34,15 @@ class AudioTranscriptionPostProcessor(PostProcessorApi):
requestor: InterProcessRequestor,
embeddings: Embeddings,
metrics: DataProcessorMetrics,
genai_manager: GenAIClientManager | None = None,
):
super().__init__(config, metrics, None)
self.config = config
self.requestor = requestor
self.embeddings = embeddings
self.genai_manager = genai_manager
self.recognizer = None
self.transcription_lock = threading.Lock()
self.transcription_thread: threading.Thread | None = None
self.transcription_running = False
self._use_genai = not isinstance(
config.audio_transcription.model, AudioTranscriptionModelEnum
)
if self._use_genai:
# never build the local recognizer on the GenAI path; WhisperModel
# downloads several hundred MB on first use
return
# faster-whisper handles model downloading automatically
self.model_path = os.path.join(MODEL_CACHE_DIR, "whisper")
@@ -163,31 +147,6 @@ class AudioTranscriptionPostProcessor(PostProcessorApi):
logger.error(f"Error in audio transcription post-processing: {e}")
def __transcribe_audio(self, audio_data: bytes) -> str | None:
"""Transcribe WAV audio data with the configured backend."""
if self._use_genai:
return self.__transcribe_audio_genai(audio_data)
return self.__transcribe_audio_whisper(audio_data)
def __transcribe_audio_genai(self, audio_data: bytes) -> str | None:
"""Hand the WAV bytes to the GenAI provider holding the transcribe role."""
client = self.genai_manager.transcribe_client if self.genai_manager else None
if not client:
logger.error(
"audio_transcription.model is '%s' (GenAI provider) but no transcribe "
"client is configured. Ensure the GenAI provider has 'transcribe' in its roles",
self.config.audio_transcription.model,
)
return None
text = client.transcribe(
audio_data,
language=resolve_language(self.config.audio_transcription.language),
)
return clean_transcript(text) or None
def __transcribe_audio_whisper(self, audio_data: bytes) -> str | None:
"""Transcribe WAV audio data using faster-whisper."""
if not self.recognizer:
logger.debug("Recognizer not initialized")
@@ -201,7 +160,7 @@ class AudioTranscriptionPostProcessor(PostProcessorApi):
segments, info = self.recognizer.transcribe(
temp_wav,
language=resolve_language(self.config.audio_transcription.language),
language=self.config.audio_transcription.language,
beam_size=5,
)
@@ -253,7 +253,7 @@ class ObjectDescriptionProcessor(PostProcessorApi):
# Crop snapshot based on region
# provide full image if region doesn't exist (manual events)
height, width = img.shape[:2]
x1_rel, y1_rel, width_rel, height_rel = event.data.get(
x1_rel, y1_rel, width_rel, height_rel = event.data.get( # type: ignore[attr-defined]
"region", [0, 0, 1, 1]
)
x1, y1 = int(x1_rel * width), int(y1_rel * height)
@@ -1,394 +0,0 @@
"""Frame annotations derived from object tracking data.
Builds short notes describing what changed during a review item, keyed to the
frames sampled from it. Everything here comes from tracked object data already
in the database (each event's `path_data` trajectory and the timeline's
stationary/active changes), so the notes can be stated to the model as fact
rather than as something it must perceive.
"""
import logging
import math
from collections.abc import Sequence
from typing import Any
from frigate.models import Event, Timeline
logger = logging.getLogger(__name__)
# Movement smaller than this (normalized frame units) between two path points
# is treated as the object holding still rather than travelling.
STILL_THRESHOLD = 0.02
# A heading change beyond this (dot product against the leg's own heading)
# counts as the object turning back rather than curving.
REVERSAL_DOT = -0.3
# A run of travel shorter than this (normalized frame units) is treated as
# milling about rather than going somewhere. Without it, a subject pacing in
# one spot produces a burst of contradictory "turns around" notes on a single
# frame.
MIN_LEG_DISTANCE = 0.08
# Movement that begins within this many seconds of detection is folded into
# the detection note, so each arrival reads as one event instead of several.
DETECT_MOVE_MERGE_SECONDS = 2.0
STATE_CHANGE_PHRASES = {
"stationary": "has stopped moving",
"active": "starts moving again",
}
Point = tuple[float, float, float]
Leg = tuple[int, int]
def describe_position(x: float, y: float) -> str:
"""Name a normalized frame position in plain terms."""
horizontal = "left" if x < 0.34 else ("right" if x > 0.66 else "center")
vertical = "top" if y < 0.34 else ("bottom" if y > 0.66 else "middle")
if horizontal == "center" and vertical == "middle":
return "the middle of the frame"
if horizontal == "center":
return f"the {vertical} of the frame"
if vertical == "middle":
return f"the {horizontal} of the frame"
return f"the {vertical} {horizontal} of the frame"
def describe_heading(dx: float, dy: float) -> str:
"""Name a direction of travel in frame terms.
y grows downward in normalized coordinates, so a falling y reads as moving
toward the top of the frame.
"""
parts = []
if abs(dy) > abs(dx) * 0.4:
parts.append("down" if dy > 0 else "up")
if abs(dx) > abs(dy) * 0.4:
parts.append("right" if dx > 0 else "left")
return " and ".join(parts) if parts else "in place"
def event_name(event: dict[str, Any]) -> str:
"""Name an object for the notes, e.g. 'a person' or 'waste bin "Compost"'.
Objects are never numbered or given track identifiers. Frigate opens a new
tracked object whenever a subject is re-detected, so the tracking data
cannot say whether two entries are the same subject, and the notes stay
ambiguous rather than implying either answer.
"""
label = str(event["label"]).replace("_", " ").replace("-verified", "")
sub_label = event.get("sub_label")
if sub_label:
return f'{label} "{sub_label}"'
article = "an" if label[:1].lower() in "aeiou" else "a"
return f"{article} {label}"
def path_legs(points: list[Point]) -> list[Leg]:
"""Split a trajectory into runs of travel in a consistent direction.
A leg ends when the subject starts moving back against the direction that
leg established, and only once the leg has covered MIN_LEG_DISTANCE, so
jitter around a standing subject does not register as a turn.
Returns (start, end) index pairs into `points`.
"""
legs: list[Leg] = []
start = 0
for i in range(1, len(points)):
lx = points[i][0] - points[start][0]
ly = points[i][1] - points[start][1]
leg_distance = (lx * lx + ly * ly) ** 0.5
if leg_distance < MIN_LEG_DISTANCE:
continue
sx = points[i][0] - points[i - 1][0]
sy = points[i][1] - points[i - 1][1]
step = (sx * sx + sy * sy) ** 0.5
if step < STILL_THRESHOLD:
continue
dot = (lx / leg_distance) * (sx / step) + (ly / leg_distance) * (sy / step)
if dot < REVERSAL_DOT:
legs.append((start, i - 1))
start = i - 1
if start < len(points) - 1:
legs.append((start, len(points) - 1))
return [
(a, b)
for a, b in legs
if ((points[b][0] - points[a][0]) ** 2 + (points[b][1] - points[a][1]) ** 2)
** 0.5
>= MIN_LEG_DISTANCE
]
def path_points(path_data: list[Any]) -> list[Point]:
"""Flatten path_data into (x, y, timestamp) tuples, or [] if malformed."""
try:
return [(p[0][0], p[0][1], p[1]) for p in path_data or []]
except (IndexError, TypeError):
logger.debug("Malformed path_data, skipping trajectory notes")
return []
def leg_start_time(points: list[Point], leg: Leg) -> float:
"""When a leg's movement actually began.
path_data always keeps an object's first two samples, so a leg can open
with points recorded long before the object moved. The first sample that
has left the leg's origin is the earliest evidence of movement.
"""
a, b = leg
x0, y0, t0 = points[a]
for x, y, t in points[a + 1 : b + 1]:
if math.hypot(x - x0, y - y0) >= STILL_THRESHOLD:
return t
return t0
def leg_heading(points: list[Point], leg: Leg) -> str:
a, b = leg
return describe_heading(points[b][0] - points[a][0], points[b][1] - points[a][1])
def leg_phrase(points: list[Point], leg: Leg, first: bool) -> str:
"""Describe the start of a leg, e.g. 'turns around at ... and heads left'."""
heading = leg_heading(points, leg)
place = describe_position(points[leg[0]][0], points[leg[0]][1])
if first:
return f"starts moving {heading} from {place}"
return f"turns around at {place} and heads {heading}"
def path_moments(path_data: list[Any]) -> list[tuple[float, str]]:
"""Key moments in one trajectory as (timestamp, phrase).
Emits one note per leg of travel. Where the last leg ends is left out:
path_data only records significant movement, so its final point cannot
distinguish an object coming to rest from one leaving the frame.
"""
points = path_points(path_data)
if len(points) < 2:
return []
return [
(leg_start_time(points, leg), leg_phrase(points, leg, index == 0))
for index, leg in enumerate(path_legs(points))
]
def build_timeline(
events: list[dict[str, Any]],
span_end: float,
state_changes: Sequence[dict[str, Any]] = (),
) -> list[tuple[float, str]]:
"""All annotated moments across every event, in time order.
Only changes are noted, since those are what sparse frames miss; an
object's state at the end of the clip is visible in the last frame.
`span_end` is the timestamp of the last sampled frame, and moments past it
describe nothing the model can see. A track ending means the object
stopped being detected, which may or may not mean it left the frame.
`state_changes` are timeline rows (timestamp, source_id, class_type); the
stationary and active ones become "has stopped moving" / "starts moving
again". Frigate only marks an object stationary after it has been still
for a while, which the past-tense wording reflects.
Each object keeps its own notes. Folding an object into the note of the
person moving it ("alongside ...") was tried and made models lose track of
where the object went.
"""
changes_by_event: dict[str, list[tuple[float, str]]] = {}
for change in state_changes:
phrase = STATE_CHANGE_PHRASES.get(change["class_type"])
if phrase:
changes_by_event.setdefault(change["source_id"], []).append(
(change["timestamp"], phrase)
)
timeline: list[tuple[float, str]] = []
for event in sorted(events, key=lambda e: e["start_time"]):
# Frame extraction can come up short at the end of a clip, leaving
# objects that only appear after the last frame we actually have.
if event["start_time"] > span_end:
continue
points = path_points(event.get("path_data") or [])
legs = path_legs(points) if len(points) >= 2 else []
name = event_name(event)
detected_at = event["start_time"]
where = describe_position(points[0][0], points[0][1]) if points else "the frame"
merges = bool(legs) and leg_start_time(points, legs[0]) - detected_at <= (
DETECT_MOVE_MERGE_SECONDS
)
remaining = list(enumerate(legs))
moments: list[tuple[float, str]] = []
if merges:
heading = leg_heading(points, legs[0])
timeline.append(
(detected_at, f"{name} first detected at {where}, moving {heading}")
)
remaining = remaining[1:]
else:
timeline.append((detected_at, f"{name} first detected at {where}"))
for index, leg in remaining:
moments.append(
(
leg_start_time(points, leg),
leg_phrase(points, leg, index == 0),
)
)
moments.extend(
(timestamp, phrase)
for timestamp, phrase in changes_by_event.get(event["id"], [])
if timestamp >= detected_at
)
for timestamp, phrase in moments:
if timestamp <= span_end:
timeline.append((timestamp, f"{name} {phrase}"))
if event["end_time"] and event["end_time"] <= span_end:
timeline.append((event["end_time"], f"{name} is no longer detected"))
return sorted(timeline, key=lambda m: m[0])
def annotations_by_frame(
timeline: list[tuple[float, str]], frame_times: list[float]
) -> dict[int, list[str]]:
"""Bucket timeline moments onto the frame that follows each one.
A moment is attached to the first frame at or after it happened, so the
note always precedes the image in which the change becomes visible.
"""
buckets: dict[int, list[str]] = {}
if not frame_times:
return buckets
for timestamp, phrase in timeline:
index = next(
(i for i, ft in enumerate(frame_times) if ft >= timestamp),
len(frame_times) - 1,
)
buckets.setdefault(index, []).append(phrase)
return buckets
def get_tracked_events(detection_ids: list[str]) -> list[dict[str, Any]]:
"""Load the tracked objects behind a review item's detections."""
if not detection_ids:
return []
rows = list(
Event.select(
Event.id,
Event.label,
Event.sub_label,
Event.start_time,
Event.end_time,
Event.data,
)
.where(Event.id << detection_ids)
.dicts()
.iterator()
)
return [
{
"id": row["id"],
"label": row["label"],
"sub_label": row["sub_label"],
"start_time": row["start_time"],
"end_time": row["end_time"],
"path_data": (row["data"] or {}).get("path_data") or [],
}
for row in rows
if row["start_time"] is not None
]
def get_state_changes(detection_ids: list[str]) -> list[dict[str, Any]]:
"""Stationary/active changes the timeline recorded for these objects."""
if not detection_ids:
return []
return list(
Timeline.select(Timeline.timestamp, Timeline.source_id, Timeline.class_type)
.where(
(Timeline.source_id << detection_ids)
& (Timeline.class_type << list(STATE_CHANGE_PHRASES))
)
.dicts()
.iterator()
)
def build_frame_captions(
detection_ids: list[str],
frame_times: list[float],
) -> list[str]:
"""A caption for each sampled frame, in frame order.
Every frame gets its index and elapsed time so the model can tell them
apart; frames where something changed also carry the tracker notes for
that moment. Returns an empty list when there is nothing to say, which
callers treat as a reason to fall back to sending plain frames.
"""
if not frame_times:
return []
events = get_tracked_events(detection_ids)
if not events:
logger.debug("No tracked events found for review item, skipping annotations")
return []
timeline = build_timeline(events, frame_times[-1], get_state_changes(detection_ids))
buckets = annotations_by_frame(timeline, frame_times)
if not buckets:
return []
total = len(frame_times)
origin = frame_times[0]
captions: list[str] = []
for index, timestamp in enumerate(frame_times):
lines = [f"Frame {index + 1} of {total} (+{timestamp - origin:.1f}s):"]
lines.extend(f"[tracker] {note}" for note in buckets.get(index, []))
captions.append("\n".join(lines))
return captions
@@ -8,7 +8,7 @@ import os
import shutil
import threading
from pathlib import Path
from typing import Any, cast
from typing import Any
import cv2
from peewee import DoesNotExist
@@ -19,11 +19,7 @@ from frigate.comms.embeddings_updater import EmbeddingsRequestEnum
from frigate.comms.inter_process import InterProcessRequestor
from frigate.config import FrigateConfig
from frigate.config.camera import CameraConfig
from frigate.config.camera.review import (
GenAIReviewConfig,
ImageSourceEnum,
ReviewFrameModeEnum,
)
from frigate.config.camera.review import GenAIReviewConfig, ImageSourceEnum
from frigate.const import (
ATTRIBUTE_LABEL_DISPLAY_MAP,
CACHE_DIR,
@@ -40,7 +36,6 @@ from frigate.util.image import get_image_from_recording
from ..post.api import PostProcessorApi
from ..types import DataProcessorMetrics
from .review_annotations import build_frame_captions
logger = logging.getLogger(__name__)
@@ -48,7 +43,6 @@ RECORDING_BUFFER_EXTENSION_PERCENT = 0.10
MIN_RECORDING_DURATION = 10
MAX_IMAGE_TOKENS = 24000
MAX_FRAMES_PER_SECOND = 1
MAX_ANNOTATED_FRAMES = 28
class ReviewDescriptionProcessor(PostProcessorApi):
@@ -73,7 +67,6 @@ class ReviewDescriptionProcessor(PostProcessorApi):
duration: float,
image_source: ImageSourceEnum = ImageSourceEnum.preview,
height: int = 480,
frame_mode: ReviewFrameModeEnum = ReviewFrameModeEnum.frames,
) -> int:
"""Calculate optimal number of frames based on event duration, context size,
image source, and resolution.
@@ -87,8 +80,6 @@ class ReviewDescriptionProcessor(PostProcessorApi):
- MAX_FRAMES_PER_SECOND x duration, to avoid drowning short events in
near-duplicate frames where the model latches onto the redundant middle
and skips the start/end action
- MAX_ANNOTATED_FRAMES in annotated mode, where the tracking notes
already carry the sequence
"""
client = self.genai_manager.description_client
@@ -134,10 +125,6 @@ class ReviewDescriptionProcessor(PostProcessorApi):
max_frames_by_tokens = int(image_token_budget / tokens_per_image)
max_frames_by_duration = int(duration * MAX_FRAMES_PER_SECOND)
max_frames = min(max_frames_by_tokens, max_frames_by_duration)
if frame_mode == ReviewFrameModeEnum.annotated_frames:
max_frames = min(max_frames, MAX_ANNOTATED_FRAMES)
return max(max_frames, 3)
def process_data(
@@ -179,7 +166,6 @@ class ReviewDescriptionProcessor(PostProcessorApi):
return
image_source = camera_config.review.genai.image_source
frame_mode = camera_config.review.genai.frame_mode
if image_source == ImageSourceEnum.recordings:
buffer_extension = get_recording_buffer_extension(
@@ -188,43 +174,40 @@ class ReviewDescriptionProcessor(PostProcessorApi):
final_data["start_time"] -= buffer_extension
final_data["end_time"] += buffer_extension
frames = self.get_recording_frames(
thumbs = self.get_recording_frames(
camera,
final_data["start_time"],
final_data["end_time"],
height=480, # Use 480p for good balance between quality and token usage
frame_mode=frame_mode,
)
if not frames:
if not thumbs:
# Fallback to preview frames if no recordings available
logger.warning(
f"No recording frames found for {camera}, falling back to preview frames"
)
frames = self.get_preview_frames_as_bytes(
thumbs = self.get_preview_frames_as_bytes(
camera,
final_data["start_time"],
final_data["end_time"],
final_data["thumb_path"],
id,
camera_config.review.genai.debug_save_thumbnails,
frame_mode,
)
elif camera_config.review.genai.debug_save_thumbnails:
self.save_debug_recording_frames(id, frames)
self.save_debug_recording_frames(id, thumbs)
else:
# Use preview frames
frames = self.get_preview_frames_as_bytes(
thumbs = self.get_preview_frames_as_bytes(
camera,
final_data["start_time"],
final_data["end_time"],
final_data["thumb_path"],
id,
camera_config.review.genai.debug_save_thumbnails,
frame_mode,
)
self.start_analysis(camera_config, final_data, frames)
self.start_analysis(camera_config, final_data, thumbs)
def handle_request(self, topic: str, request_data: dict[str, Any]) -> str | None:
if topic == EmbeddingsRequestEnum.regenerate_review_description.value:
@@ -403,50 +386,31 @@ class ReviewDescriptionProcessor(PostProcessorApi):
buffer_extension = get_recording_buffer_extension(
final_data["end_time"] - final_data["start_time"]
)
frames = self.get_recording_frames(
thumbs = self.get_recording_frames(
str(review.camera),
final_data["start_time"] - buffer_extension,
final_data["end_time"] + buffer_extension,
height=480,
frame_mode=camera_config.review.genai.frame_mode,
)
if not frames:
if not thumbs:
logger.error(
"No recording frames are available for review item %s", review_id
)
return
if camera_config.review.genai.debug_save_thumbnails:
self.save_debug_recording_frames(review_id, frames)
self.save_debug_recording_frames(review_id, thumbs)
self.start_analysis(camera_config, final_data, frames)
self.start_analysis(camera_config, final_data, thumbs)
def start_analysis(
self,
camera_config: CameraConfig,
final_data: dict[str, Any],
frames: list[tuple[bytes, float]],
thumbs: list[bytes],
) -> None:
"""Kick off description generation for a review item in the background."""
thumbs = [frame for frame, _ in frames]
captions: list[str] = []
if (
camera_config.review.genai.frame_mode
== ReviewFrameModeEnum.annotated_frames
):
captions = build_frame_captions(
final_data["data"].get("detections") or [],
[timestamp for _, timestamp in frames],
)
if not captions:
logger.debug(
"No tracking annotations for review item %s, sending plain frames",
final_data["id"],
)
self.review_desc_dps.update()
threading.Thread(
target=run_analysis,
@@ -457,22 +421,19 @@ class ReviewDescriptionProcessor(PostProcessorApi):
camera_config,
final_data,
thumbs,
captions,
camera_config.review.genai,
sorted(self.config.all_labels),
self.config.all_attributes,
),
).start()
def save_debug_recording_frames(
self, review_id: str, frames: list[tuple[bytes, float]]
) -> None:
def save_debug_recording_frames(self, review_id: str, thumbs: list[bytes]) -> None:
"""Write the recording frames sent to the provider out for debugging."""
Path(os.path.join(CLIPS_DIR, "genai-requests", review_id)).mkdir(
parents=True, exist_ok=True
)
for idx, (frame_bytes, _) in enumerate(frames):
for idx, frame_bytes in enumerate(thumbs):
with open(
os.path.join(CLIPS_DIR, f"genai-requests/{review_id}/{idx}.jpg"),
"wb",
@@ -484,9 +445,7 @@ class ReviewDescriptionProcessor(PostProcessorApi):
camera: str,
start_time: float,
end_time: float,
frame_mode: ReviewFrameModeEnum = ReviewFrameModeEnum.frames,
) -> list[tuple[str, float]]:
"""Preview frame paths paired with the time each one was captured."""
) -> list[str]:
preview_dir = os.path.join(CACHE_DIR, "preview_frames")
file_start = f"preview_{camera}-"
start_file = f"{file_start}{start_time}.webp"
@@ -518,29 +477,18 @@ class ReviewDescriptionProcessor(PostProcessorApi):
frame_count = len(all_frames)
desired_frame_count = self.calculate_frame_count(
camera,
duration=end_time - start_time,
frame_mode=frame_mode,
camera, duration=end_time - start_time
)
def with_timestamp(path: str) -> tuple[str, float]:
# Preview frames are named preview_<camera>-<timestamp>.webp
stem = os.path.basename(path).removesuffix(".webp")
try:
return (path, float(stem.removeprefix(file_start)))
except ValueError:
return (path, start_time)
if frame_count <= desired_frame_count:
return [with_timestamp(f) for f in all_frames]
return all_frames
selected_frames = []
step_size = (frame_count - 1) / (desired_frame_count - 1)
for i in range(desired_frame_count):
index = round(i * step_size)
selected_frames.append(with_timestamp(all_frames[index]))
selected_frames.append(all_frames[index])
return selected_frames
@@ -550,12 +498,11 @@ class ReviewDescriptionProcessor(PostProcessorApi):
start_time: float,
end_time: float,
height: int = 480,
frame_mode: ReviewFrameModeEnum = ReviewFrameModeEnum.frames,
) -> list[tuple[bytes, float]]:
"""Get frames from recordings paired with the time each was captured."""
) -> list[bytes]:
"""Get frames from recordings at specified timestamps."""
duration = end_time - start_time
desired_frame_count = self.calculate_frame_count(
camera, duration, ImageSourceEnum.recordings, height, frame_mode
camera, duration, ImageSourceEnum.recordings, height
)
# Calculate evenly spaced timestamps throughout the duration
@@ -581,8 +528,7 @@ class ReviewDescriptionProcessor(PostProcessorApi):
.get()
)
# start_time is a DateTimeField holding a unix timestamp
time_in_segment = ts - cast(float, recording.start_time)
time_in_segment = ts - recording.start_time
return get_image_from_recording(
self.config.ffmpeg,
recording.path,
@@ -593,7 +539,7 @@ class ReviewDescriptionProcessor(PostProcessorApi):
except DoesNotExist:
return None
frames: list[tuple[bytes, float]] = []
frames = []
for timestamp in timestamps:
try:
@@ -606,7 +552,7 @@ class ReviewDescriptionProcessor(PostProcessorApi):
image_data = extract_frame_from_recording(rounded_timestamp)
if image_data:
frames.append((image_data, timestamp))
frames.append(image_data)
else:
logger.warning(
f"No recording found for {camera} at timestamp {timestamp}"
@@ -627,8 +573,7 @@ class ReviewDescriptionProcessor(PostProcessorApi):
thumb_path_fallback: str,
review_id: str,
save_debug: bool,
frame_mode: ReviewFrameModeEnum = ReviewFrameModeEnum.frames,
) -> list[tuple[bytes, float]]:
) -> list[bytes]:
"""Get preview frames and convert them to JPEG bytes.
Args:
@@ -640,14 +585,14 @@ class ReviewDescriptionProcessor(PostProcessorApi):
save_debug: Whether to save debug thumbnails
Returns:
List of (JPEG image bytes, capture timestamp) pairs
List of JPEG image bytes
"""
frame_paths = self.get_cache_frames(camera, start_time, end_time, frame_mode)
frame_paths = self.get_cache_frames(camera, start_time, end_time)
if not frame_paths:
frame_paths = [(thumb_path_fallback, start_time)]
frame_paths = [thumb_path_fallback]
thumbs: list[tuple[bytes, float]] = []
for idx, (thumb_path, timestamp) in enumerate(frame_paths):
thumbs = []
for idx, thumb_path in enumerate(frame_paths):
thumb_data = cv2.imread(thumb_path)
if thumb_data is None:
@@ -660,7 +605,7 @@ class ReviewDescriptionProcessor(PostProcessorApi):
".jpg", thumb_data, [int(cv2.IMWRITE_JPEG_QUALITY), 100]
)
if ret:
thumbs.append((jpg.tobytes(), timestamp))
thumbs.append(jpg.tobytes())
if save_debug:
Path(os.path.join(CLIPS_DIR, "genai-requests", review_id)).mkdir(
@@ -695,7 +640,6 @@ def run_analysis(
camera_config: CameraConfig,
final_data: dict[str, Any],
thumbs: list[bytes],
frame_captions: list[str],
genai_config: GenAIReviewConfig,
labelmap_objects: list[str],
attribute_labels: list[str],
@@ -752,7 +696,6 @@ def run_analysis(
genai_config.debug_save_thumbnails,
genai_config.activity_context_prompt,
genai_config.response_style,
frame_captions,
)
review_inference_speed.update(datetime.datetime.now().timestamp() - start)
@@ -237,7 +237,7 @@ class SemanticTriggerProcessor(PostProcessorApi):
return
# Skip the event if not an object
if event.data.get("type") != "object":
if event.data.get("type") != "object": # type: ignore[attr-defined]
return
thumbnail_bytes = get_event_thumbnail_bytes(event)
@@ -1,19 +1,16 @@
"""Handle processing audio for speech transcription using sherpa-onnx with FFmpeg pipe."""
import collections
import logging
import os
import queue
import threading
import time
from typing import TYPE_CHECKING, Any
from typing import Any
import numpy as np
from frigate.comms.inter_process import InterProcessRequestor
from frigate.config import CameraConfig, FrigateConfig
from frigate.config.classification import AudioTranscriptionModelEnum
from frigate.const import AUDIO_DURATION, MODEL_CACHE_DIR
from frigate.const import MODEL_CACHE_DIR
from frigate.data_processing.common.audio_transcription.model import (
AudioTranscriptionModelRunner,
)
@@ -21,38 +18,12 @@ from frigate.data_processing.real_time.whisper_online import (
FasterWhisperASR,
OnlineASRProcessor,
)
from frigate.util.audio import (
clean_transcript,
pcm16_to_wav,
resolve_language,
stitch_transcripts,
)
from ..types import DataProcessorMetrics
from .api import RealTimeProcessorApi
if TYPE_CHECKING:
# importing frigate.genai eagerly would pull the provider SDKs into the
# audio process even when transcription runs on a local model
from frigate.genai.manager import GenAIClientManager
logger = logging.getLogger(__name__)
# Number of ~0.975s audio detector chunks per GenAI request. The window advances
# one chunk at a time, so two chunks means a 50% overlap: every word lands whole
# in at least one window, which whisper-family models need to avoid hallucinating
# on a clipped clip. The cadence is fixed by the audio detector's frame size, so
# this is a constant rather than a config knob.
GENAI_WINDOW_CHUNKS = 2
# Bound the queue at ~30s of audio so a slow or hung provider cannot grow it
# without limit. The producer is the ffmpeg read thread and must never block.
AUDIO_QUEUE_MAXSIZE = int(30 / AUDIO_DURATION)
# A backed-up queue drops a chunk per cycle, so warning on each one would spam
# the log once a second per camera for as long as the provider stays slow.
AUDIO_DROP_WARN_INTERVAL = 10.0
class AudioTranscriptionRealTimeProcessor(RealTimeProcessorApi):
def __init__(
@@ -60,10 +31,9 @@ class AudioTranscriptionRealTimeProcessor(RealTimeProcessorApi):
config: FrigateConfig,
camera_config: CameraConfig,
requestor: InterProcessRequestor,
model_runner: AudioTranscriptionModelRunner | None,
model_runner: AudioTranscriptionModelRunner,
metrics: DataProcessorMetrics,
stop_event: threading.Event,
genai_manager: "GenAIClientManager | None" = None,
):
super().__init__(config, metrics)
self.config = config
@@ -72,31 +42,11 @@ class AudioTranscriptionRealTimeProcessor(RealTimeProcessorApi):
self.stream: Any = None
self.whisper_model: FasterWhisperASR | None = None
self.model_runner = model_runner
self.genai_manager = genai_manager
self.transcription_segments: list[str] = []
self.audio_queue: queue.Queue[tuple[dict[str, Any], np.ndarray]] = queue.Queue(
maxsize=AUDIO_QUEUE_MAXSIZE
)
self.audio_queue: queue.Queue[tuple[dict[str, Any], np.ndarray]] = queue.Queue()
self.stop_event = stop_event
self._use_genai = not isinstance(
config.audio_transcription.model, AudioTranscriptionModelEnum
)
# sliding window of raw int16 chunks; the deque's maxlen is what evicts
# the oldest chunk and so produces the overlap
self._genai_window: collections.deque[np.ndarray] = collections.deque(
maxlen=GENAI_WINDOW_CHUNKS
)
self._genai_committed = ""
# set by the producer when it discards a chunk, so the consumer knows the
# audio it is about to receive is not contiguous with what it buffered
self._audio_dropped = threading.Event()
self._last_drop_warning = 0.0
def __build_recognizer(self) -> None:
if self._use_genai:
# nothing local to load; never import sherpa or FasterWhisperASR
return
try:
if self.config.audio_transcription.model_size == "large":
# Whisper models need to be per-process and can only run one stream at a time
@@ -114,7 +64,7 @@ class AudioTranscriptionRealTimeProcessor(RealTimeProcessorApi):
self.stream = OnlineASRProcessor(
asr=self.whisper_model,
)
elif self.model_runner is not None:
else:
logger.debug(f"Loading sherpa stream for {self.camera_config.name}")
self.stream = self.model_runner.model.create_stream()
logger.debug(
@@ -126,15 +76,6 @@ class AudioTranscriptionRealTimeProcessor(RealTimeProcessorApi):
)
def __process_audio_stream(self, audio_data: np.ndarray) -> tuple[str, bool] | None:
# must precede both the model_runner guard (model_runner is None on this
# path) and the float32 normalization below (GenAI wants untouched int16)
if self._use_genai:
return self.__process_audio_genai(audio_data)
if self.model_runner is None:
logger.debug("Audio transcription (live) model runner not initialized")
return None
if (
self.model_runner.model is None
and self.config.audio_transcription.model_size == "small"
@@ -199,82 +140,6 @@ class AudioTranscriptionRealTimeProcessor(RealTimeProcessorApi):
logger.error(f"Error processing audio stream: {e}")
return None
def __process_audio_genai(self, audio_data: np.ndarray) -> tuple[str, bool] | None:
"""Transcribe a sliding overlapped window through the GenAI provider."""
client = self.genai_manager.transcribe_client if self.genai_manager else None
if not client:
logger.error(
"audio_transcription.model is '%s' (GenAI provider) but no transcribe "
"client is configured. Ensure the GenAI provider has 'transcribe' in its roles",
self.config.audio_transcription.model,
)
return None
if self._audio_dropped.is_set():
self._audio_dropped.clear()
# Chunks were discarded between what is buffered and this one, so
# concatenating them would splice non-adjacent audio into one window
# and destroy the overlap the stitcher depends on.
self._genai_window.clear()
if self._genai_committed:
# the transcript has a gap in it; close the utterance out rather
# than stitching across missing speech
return self.__end_genai_utterance()
self._genai_window.append(audio_data)
if len(self._genai_window) < GENAI_WINDOW_CHUNKS:
# wait for a full window so the first request is never a clipped clip
return None
window = np.concatenate(list(self._genai_window))
# Silence gate, using the same threshold audio detection uses. Gate the
# whole window rather than individual chunks; this is the primary cost
# and privacy brake and is what keeps a quiet camera near zero requests.
window_as_float = window.astype(np.float32)
rms = float(np.sqrt(np.mean(np.absolute(np.square(window_as_float)))))
if rms < self.camera_config.audio.min_volume:
logger.debug(
f"Window RMS {rms:.1f} below min_volume, skipping transcription"
)
return self.__end_genai_utterance()
text = client.transcribe(
pcm16_to_wav(window),
language=resolve_language(self.config.audio_transcription.language),
)
# cleaning has to come first: a silent window often comes back as the
# model's preamble alone, which is silence, not a word to commit
cleaned = clean_transcript(text)
if not cleaned:
return self.__end_genai_utterance()
self._genai_committed = stitch_transcripts(self._genai_committed, cleaned)
# no VAD on this path, so mirror the whisper branch's heuristic endpoint
is_endpoint = (
self._genai_committed.endswith((".", "!", "?"))
and len(self._genai_committed) > 300
)
logger.debug(f"GenAI transcription: '{self._genai_committed}'")
return self._genai_committed, is_endpoint
def __end_genai_utterance(self) -> tuple[str, bool] | None:
"""Close out the current utterance when a window carries no speech."""
if not self._genai_committed:
return None
return self._genai_committed, True
def process_frame(self, obj_data: dict[str, Any], frame: np.ndarray) -> None:
pass
@@ -283,38 +148,8 @@ class AudioTranscriptionRealTimeProcessor(RealTimeProcessorApi):
logger.debug("No audio data provided for transcription")
return None
# enqueue audio data for processing in the thread. never block: the
# producer is the ffmpeg read thread that audio detection depends on,
# so on a backlog drop the oldest chunk instead.
try:
self.audio_queue.put_nowait((obj_data, audio))
except queue.Full:
try:
self.audio_queue.get_nowait()
self.audio_queue.task_done()
except queue.Empty:
pass
# the stream now has a hole in it, which the consumer has to know
# about before it splices the next chunk onto what it already holds
self._audio_dropped.set()
now = time.monotonic()
if now - self._last_drop_warning >= AUDIO_DROP_WARN_INTERVAL:
self._last_drop_warning = now
logger.warning(
"Audio transcription queue for %s is full, dropping audio. The "
"provider is not keeping up with the %.2fs chunk rate",
self.camera_config.name,
AUDIO_DURATION,
)
try:
self.audio_queue.put_nowait((obj_data, audio))
except queue.Full:
pass
# enqueue audio data for processing in the thread
self.audio_queue.put((obj_data, audio))
return None
def run(self) -> None:
@@ -370,14 +205,6 @@ class AudioTranscriptionRealTimeProcessor(RealTimeProcessorApi):
break
def reset(self) -> None:
if self._use_genai:
self._genai_committed = ""
# stale audio carried across an utterance boundary would be
# re-transcribed into the next one
self._genai_window.clear()
logger.debug("Stream reset")
return
if self.config.audio_transcription.model_size == "large":
# get final output from whisper
output = self.stream.finish()
@@ -391,7 +218,7 @@ class AudioTranscriptionRealTimeProcessor(RealTimeProcessorApi):
# reset whisper
self.stream.init()
self.transcription_segments = []
elif self.model_runner is not None:
else:
# reset sherpa
self.model_runner.model.reset(self.stream)
@@ -399,24 +226,6 @@ class AudioTranscriptionRealTimeProcessor(RealTimeProcessorApi):
def check_unload_model(self) -> None:
# regularly called in the loop in audio maintainer
if self._use_genai:
# no model to unload, but this is the hook that fires when
# live_enabled flips off. guard on emptiness: called ~1x/s per camera.
if self._genai_committed or self._genai_window:
logger.debug(
f"Clearing GenAI transcription state for {self.camera_config.name}"
)
self.clear_audio_queue()
self._genai_committed = ""
self._genai_window.clear()
self.requestor.send_data(
f"{self.camera_config.name}/audio/transcription",
"",
)
return
if (
self.config.audio_transcription.model_size == "large"
and self.whisper_model is not None
@@ -461,10 +270,6 @@ class AudioTranscriptionRealTimeProcessor(RealTimeProcessorApi):
self, topic: str, request_data: dict[str, Any]
) -> dict[str, Any] | None:
if topic == "clear_audio_recognizer":
if self._use_genai:
self.reset()
return {"message": "Audio transcription state cleared", "success": True}
self.stream = None
self.__build_recognizer()
return {"message": "Audio recognizer cleared and rebuilt", "success": True}
+26 -37
View File
@@ -247,20 +247,15 @@ class CudaGraphRunner(BaseModelRunner):
EnrichmentModelTypeEnum.yolov9_license_plate.value,
]
# ORT performs two regular runs before it starts capturing, but on some
# driver / cuDNN combinations the arena still has to extend on the run that
# captures, and cudaMalloc is not allowed during capture. Running with
# capture disabled first keeps those allocations outside of the capture.
GRAPH_FREE_WARMUP_RUNS = 2
def __init__(self, session: ort.InferenceSession, cuda_device_id: int):
self._session = session
self._cuda_device_id = cuda_device_id
self._prepared = False
self._captured = False
self._io_binding: ort.IOBinding | None = None
self._input_name: str | None = None
self._output_names: list[str] | None = None
self._input_ortvalue: ort.OrtValue | None = None
self._output_ortvalues: ort.OrtValue | None = None
def get_input_names(self) -> list[str]:
"""Get input names for the model."""
@@ -270,41 +265,35 @@ class CudaGraphRunner(BaseModelRunner):
"""Get the input width of the model."""
return self._session.get_inputs()[0].shape[3]
def _prepare(self, input_name: str, tensor_input: np.ndarray) -> None:
"""Bind CUDA buffers and warm the session up with capture disabled."""
self._io_binding = self._session.io_binding()
self._input_name = input_name
self._output_names = [o.name for o in self._session.get_outputs()]
self._input_ortvalue = ort.OrtValue.ortvalue_from_numpy(
tensor_input, "cuda", self._cuda_device_id
)
self._io_binding.bind_ortvalue_input(self._input_name, self._input_ortvalue)
for name in self._output_names:
# Bind outputs to CUDA and allow ORT to allocate appropriately
self._io_binding.bind_output(name, "cuda", self._cuda_device_id)
# gpu_graph_id -1 disables capture and replay for the run
warmup_options = ort.RunOptions()
warmup_options.add_run_config_entry("gpu_graph_id", "-1")
for _ in range(self.GRAPH_FREE_WARMUP_RUNS):
self._session.run_with_iobinding(self._io_binding, warmup_options)
self._prepared = True
def run(self, input: dict[str, Any]):
# Extract the single tensor input (assuming one input)
input_name = list(input.keys())[0]
tensor_input = np.ascontiguousarray(input[input_name])
tensor_input = input[input_name]
tensor_input = np.ascontiguousarray(tensor_input)
if not self._prepared:
self._prepare(input_name, tensor_input)
else:
# Replay using updated input
self._input_ortvalue.update_inplace(tensor_input)
if not self._captured:
# Prepare IOBinding with CUDA buffers and let ORT allocate outputs on device
self._io_binding = self._session.io_binding()
self._input_name = input_name
self._output_names = [o.name for o in self._session.get_outputs()]
self._input_ortvalue = ort.OrtValue.ortvalue_from_numpy(
tensor_input, "cuda", self._cuda_device_id
)
self._io_binding.bind_ortvalue_input(self._input_name, self._input_ortvalue)
for name in self._output_names:
# Bind outputs to CUDA and allow ORT to allocate appropriately
self._io_binding.bind_output(name, "cuda", self._cuda_device_id)
# First IOBinding run to allocate, execute, and capture CUDA Graph
ro = ort.RunOptions()
self._session.run_with_iobinding(self._io_binding, ro)
self._captured = True
return self._io_binding.copy_outputs_to_cpu()
# Replay using updated input, copy results to CPU
self._input_ortvalue.update_inplace(tensor_input)
ro = ort.RunOptions()
self._session.run_with_iobinding(self._io_binding, ro)
return self._io_binding.copy_outputs_to_cpu()
+7 -8
View File
@@ -3,14 +3,14 @@ import json
import logging
import os
from enum import Enum
from typing import ClassVar
from typing import Any, ClassVar
import requests
from pydantic import BaseModel, ConfigDict, Field
from pydantic.fields import PrivateAttr
from frigate.const import DEFAULT_ATTRIBUTE_LABEL_MAP, MODEL_CACHE_DIR
from frigate.plus import PlusApi, load_plus_model_info
from frigate.plus import PlusApi
from frigate.util.builtin import generate_color_palette, load_labels
logger = logging.getLogger(__name__)
@@ -190,13 +190,12 @@ class ModelConfig(BaseModel):
# download the model info if it doesn't exist
if not os.path.isfile(model_info_path):
model_info = plus_api.get_model_info(model_id)
with open(model_info_path, "w") as f:
json.dump(plus_api.get_model_info(model_id), f)
model_info = load_plus_model_info(model_id)
if model_info is None:
raise ValueError(f"Unable to read the model info for {model_id}")
json.dump(model_info, f)
else:
with open(model_info_path) as f:
model_info: dict[str, Any] = json.load(f)
if detector and detector not in model_info["supportedDetectors"]:
raise ValueError(f"Model does not support detector type of {detector}")
+2 -2
View File
@@ -259,8 +259,8 @@ def detect_hailo() -> DetectionHardware | None:
# the hailo runtime schedules across every attached device itself, so there
# is nothing to address individually
units = [HardwareUnit(device="hailo:PCIe", label=os.path.basename(nodes[0]))]
return _hardware("hailo", "hailo", "Hailo", units)
units = [HardwareUnit(device="hailo8l:PCIe", label=os.path.basename(nodes[0]))]
return _hardware("hailo8l", "hailo8l", "Hailo", units)
def detect_memryx() -> DetectionHardware | None:
+105
View File
@@ -0,0 +1,105 @@
import io
import logging
from typing import Literal
import numpy as np
import requests
from PIL import Image
from pydantic import ConfigDict, Field
from frigate.detectors.detection_api import DetectionApi
from frigate.detectors.detector_config import BaseDetectorConfig
logger = logging.getLogger(__name__)
DETECTOR_KEY = "deepstack"
class DeepstackDetectorConfig(BaseDetectorConfig):
"""DeepStack/CodeProject.AI detector that sends images to a remote DeepStack HTTP API for inference. Not recommended."""
model_config = ConfigDict(
title="DeepStack",
)
type: Literal[DETECTOR_KEY]
api_url: str = Field(
default="http://localhost:80/v1/vision/detection",
title="DeepStack API URL",
description="The URL of the DeepStack API.",
)
api_timeout: float = Field(
default=0.1,
title="DeepStack API timeout (in seconds)",
description="Maximum time allowed for a DeepStack API request.",
)
api_key: str = Field(
default="",
title="DeepStack API key (if required)",
description="Optional API key for authenticated DeepStack services.",
)
class DeepStack(DetectionApi):
type_key = DETECTOR_KEY
def __init__(self, detector_config: DeepstackDetectorConfig):
self.api_url = detector_config.api_url
self.api_timeout = detector_config.api_timeout
self.api_key = detector_config.api_key
self.labels = detector_config.model.merged_labelmap
self.session = requests.Session()
def get_label_index(self, label_value):
if label_value.lower() == "truck":
label_value = "car"
for index, value in self.labels.items():
if value == label_value.lower():
return index
return -1
def detect_raw(self, tensor_input):
image_data = np.squeeze(tensor_input).astype(np.uint8)
image = Image.fromarray(image_data)
self.w, self.h = image.size
with io.BytesIO() as output:
image.save(output, format="JPEG")
image_bytes = output.getvalue()
data = {"api_key": self.api_key}
try:
response = self.session.post(
self.api_url,
data=data,
files={"image": image_bytes},
timeout=self.api_timeout,
)
except requests.exceptions.RequestException as ex:
logger.error("Error calling deepstack API: %s", ex)
return np.zeros((20, 6), np.float32)
response_json = response.json()
detections = np.zeros((20, 6), np.float32)
if response_json.get("predictions") is None:
logger.debug(f"Error in parsing response json: {response_json}")
return detections
for i, detection in enumerate(response_json.get("predictions")):
logger.debug(f"Response: {detection}")
if detection["confidence"] < 0.4:
logger.debug("Break due to confidence < 0.4")
break
label = self.get_label_index(detection["label"])
if label < 0:
logger.debug("Break due to unknown label")
break
detections[i] = [
label,
float(detection["confidence"]),
detection["y_min"] / self.h,
detection["x_min"] / self.w,
detection["y_max"] / self.h,
detection["x_max"] / self.w,
]
return detections
@@ -53,7 +53,7 @@ def preprocess_tensor(image: np.ndarray, model_w: int, model_h: int) -> np.ndarr
# ----------------- Global Constants ----------------- #
DETECTOR_KEY = "hailo"
DETECTOR_KEY = "hailo8l"
ARCH = None
H8_DEFAULT_MODEL = "yolov6n.hef"
H8L_DEFAULT_MODEL = "yolov6n.hef"
@@ -469,10 +469,10 @@ class HailoDetector(DetectionApi):
# ----------------- HailoDetectorConfig Class ----------------- #
class HailoDetectorConfig(BaseDetectorConfig):
"""Hailo detector using HEF models and the HailoRT SDK for inference on Hailo hardware."""
"""Hailo-8/Hailo-8L detector using HEF models and the HailoRT SDK for inference on Hailo hardware."""
model_config = ConfigDict(
title="Hailo",
title="Hailo-8/Hailo-8L",
)
type: Literal[DETECTOR_KEY]
+3 -7
View File
@@ -68,7 +68,7 @@ from frigate.events.types import (
RegenerateDescriptionEnum,
)
from frigate.genai import GenAIClientManager
from frigate.models import Event, Recordings, ReviewSegment, Timeline, Trigger
from frigate.models import Event, Recordings, ReviewSegment, Trigger
from frigate.types import TrackedObjectUpdateTypesEnum
from frigate.util.builtin import serialize
from frigate.util.file import get_event_thumbnail_bytes
@@ -139,7 +139,7 @@ class EmbeddingMaintainer(threading.Thread):
),
load_vec_extension=True,
)
models = [Event, Recordings, ReviewSegment, Timeline, Trigger]
models = [Event, Recordings, ReviewSegment, Trigger]
db.bind(models)
self.genai_manager = GenAIClientManager(config)
@@ -251,11 +251,7 @@ class EmbeddingMaintainer(threading.Thread):
):
self.post_processors.append(
AudioTranscriptionPostProcessor(
self.config,
self.requestor,
self.embeddings,
metrics,
self.genai_manager,
self.config, self.requestor, self.embeddings, metrics
)
)
+38 -71
View File
@@ -11,7 +11,6 @@ from typing import Any
import numpy as np
from frigate.camera import CameraMetrics
from frigate.comms.detections_updater import DetectionPublisher, DetectionTypeEnum
from frigate.comms.inter_process import InterProcessRequestor
from frigate.config import CameraConfig, CameraInput, FrigateConfig
@@ -20,7 +19,6 @@ from frigate.config.camera.updater import (
CameraConfigUpdateEnum,
CameraConfigUpdateSubscriber,
)
from frigate.config.classification import AudioTranscriptionModelEnum
from frigate.const import (
AUDIO_DURATION,
AUDIO_FORMAT,
@@ -37,7 +35,6 @@ from frigate.data_processing.common.audio_transcription.model import (
from frigate.data_processing.real_time.audio_transcription import (
AudioTranscriptionRealTimeProcessor,
)
from frigate.data_processing.types import DataProcessorMetrics
from frigate.ffmpeg_presets import parse_preset_input
from frigate.log import LogPipe, suppress_stderr_during
from frigate.util.builtin import get_ffmpeg_arg_list, load_labels
@@ -88,7 +85,6 @@ class AudioProcessor(FrigateProcess):
self,
config: FrigateConfig,
camera_metrics: DictProxy,
embeddings_metrics: DataProcessorMetrics,
stop_event: MpEvent,
):
super().__init__(
@@ -96,38 +92,8 @@ class AudioProcessor(FrigateProcess):
)
self.camera_metrics = camera_metrics
self.embeddings_metrics = embeddings_metrics
self.config = config
def spawn_if_needed(self, camera: CameraConfig) -> None:
"""Start an audio maintainer for the camera once everything it needs
has arrived. Returning early leaves the camera for the next poll."""
name = camera.name
if name is None or name in self.audio_threads:
return
if not camera.enabled or not camera.audio.enabled:
return
# ffmpeg update may not have arrived yet
if not any("audio" in i.roles for i in camera.ffmpeg.inputs):
return
# the camera maintainer creates metrics on its own poll of the same
# add update and may not have gotten there yet
metrics = self.camera_metrics.get(name)
if metrics is None:
return
thread = AudioEventMaintainer(
camera,
self.config,
metrics,
self.embeddings_metrics,
self.transcription_model_runner,
self.stop_event, # type: ignore[arg-type]
self.genai_manager,
)
self.audio_threads[name] = thread
thread.start()
self.logger.info(f"Audio maintainer started for {name}")
def __stop_audio_thread(self, camera: str) -> None:
thread = self.audio_threads.pop(camera, None)
if thread is None:
@@ -146,31 +112,18 @@ class AudioProcessor(FrigateProcess):
threading.current_thread().name = "process:audio_manager"
self.transcription_model_runner: AudioTranscriptionModelRunner | None = None
self.genai_manager: Any = None
if any(
c.enabled_in_config and c.audio_transcription.enabled
for c in self.config.cameras.values()
):
if isinstance(
self.config.audio_transcription.model, AudioTranscriptionModelEnum
):
# AudioTranscriptionModelRunner.__init__ unconditionally fetches
# sherpa-onnx or whisper weights, so only build it on the local path
self.transcription_model_runner = AudioTranscriptionModelRunner(
self.transcription_model_runner: AudioTranscriptionModelRunner | None = (
AudioTranscriptionModelRunner(
self.config.audio_transcription.device or "AUTO",
self.config.audio_transcription.model_size,
)
else:
# imported here rather than at module scope: frigate.genai pulls in
# numpy, the provider SDKs, frigate.models, and the prompt builders,
# and this process runs at PROCESS_PRIORITY_HIGH. built after the
# fork because SDK clients hold sockets and TLS state that must not
# cross it; clients themselves stay lazy behind the role property.
from frigate.genai.manager import GenAIClientManager
self.genai_manager = GenAIClientManager(self.config)
)
else:
self.transcription_model_runner = None
config_subscriber = CameraConfigUpdateSubscriber(
self.config,
@@ -183,8 +136,28 @@ class AudioProcessor(FrigateProcess):
],
)
def spawn_if_needed(camera: CameraConfig) -> None:
name = camera.name
if name is None or name in self.audio_threads:
return
if not camera.enabled or not camera.audio.enabled:
return
# ffmpeg update may not have arrived yet; wait for next poll
if not any("audio" in i.roles for i in camera.ffmpeg.inputs):
return
thread = AudioEventMaintainer(
camera,
self.config,
self.camera_metrics,
self.transcription_model_runner,
self.stop_event, # type: ignore[arg-type]
)
self.audio_threads[name] = thread
thread.start()
self.logger.info(f"Audio maintainer started for {name}")
for camera in self.config.cameras.values():
self.spawn_if_needed(camera)
spawn_if_needed(camera)
self.logger.info(f"Audio processor started (pid: {self.pid})")
@@ -194,14 +167,15 @@ class AudioProcessor(FrigateProcess):
updated_topics = config_subscriber.check_for_updates()
# stop maintainers for removed cameras so their ffmpeg process is
# torn down
# torn down and they stop touching camera_metrics (which the camera
# maintainer has already popped for the removed camera)
for removed_camera in updated_topics.get(
CameraConfigUpdateEnum.remove.name, []
):
self.__stop_audio_thread(removed_camera)
for camera in self.config.cameras.values():
self.spawn_if_needed(camera)
spawn_if_needed(camera)
config_subscriber.stop()
@@ -223,21 +197,15 @@ class AudioEventMaintainer(threading.Thread):
self,
camera: CameraConfig,
config: FrigateConfig,
metrics: CameraMetrics,
embeddings_metrics: DataProcessorMetrics,
camera_metrics: DictProxy,
audio_transcription_model_runner: AudioTranscriptionModelRunner | None,
stop_event: threading.Event,
genai_manager: Any = None,
) -> None:
super().__init__(name=f"{camera.name}_audio_event_processor")
self.config = config
self.camera_config = camera
# hold the metrics object rather than indexing the manager dict per
# chunk, which costs an IPC round trip and breaks once the camera
# maintainer pops the entry on removal
self.metrics = metrics
self.embeddings_metrics = embeddings_metrics
self.camera_metrics = camera_metrics
self.stop_event = stop_event
# per-camera stop signal so a single maintainer can be torn down at
# runtime (e.g. on camera removal) without stopping the whole process
@@ -254,7 +222,6 @@ class AudioEventMaintainer(threading.Thread):
self.logpipe = LogPipe(f"ffmpeg.{self.camera_config.name}.audio")
self.audio_listener: subprocess.Popen[Any] | None = None
self.audio_transcription_model_runner = audio_transcription_model_runner
self.genai_manager = genai_manager
self.transcription_processor = None
self.transcription_thread = None
@@ -271,9 +238,9 @@ class AudioEventMaintainer(threading.Thread):
)
self.detection_publisher = DetectionPublisher(DetectionTypeEnum.audio.value)
if self.camera_config.audio_transcription.enabled and (
self.audio_transcription_model_runner is not None
or self.genai_manager is not None
if (
self.camera_config.audio_transcription.enabled
and self.audio_transcription_model_runner is not None
):
# init the transcription processor for this camera
self.transcription_processor = AudioTranscriptionRealTimeProcessor(
@@ -281,9 +248,8 @@ class AudioEventMaintainer(threading.Thread):
camera_config=self.camera_config,
requestor=self.requestor,
model_runner=self.audio_transcription_model_runner,
metrics=self.embeddings_metrics,
metrics=self.camera_metrics[self.camera_config.name],
stop_event=self.stop_event,
genai_manager=self.genai_manager,
)
self.transcription_thread = threading.Thread(
@@ -307,8 +273,8 @@ class AudioEventMaintainer(threading.Thread):
audio_as_float: np.ndarray = audio.astype(np.float32)
rms, dBFS = self.calculate_audio_levels(audio_as_float)
self.metrics.audio_rms.value = rms
self.metrics.audio_dBFS.value = dBFS
self.camera_metrics[self.camera_config.name].audio_rms.value = rms
self.camera_metrics[self.camera_config.name].audio_dBFS.value = dBFS
audio_detections: list[tuple[str, float]] = []
@@ -393,6 +359,7 @@ class AudioEventMaintainer(threading.Thread):
return
time.sleep(self.camera_config.ffmpeg.retry_interval)
self.logpipe.dump()
self.start_or_restart_ffmpeg()
if self.audio_listener is None or self.audio_listener.stdout is None:
+5 -5
View File
@@ -37,7 +37,7 @@ class EventCleanup(threading.Thread):
if self.removed_camera_labels is None:
self.removed_camera_labels = list(
Event.select(Event.label)
.where(Event.camera.not_in(self.camera_keys))
.where(Event.camera.not_in(self.camera_keys)) # type: ignore[arg-type,call-arg,misc]
.distinct()
.execute()
)
@@ -89,7 +89,7 @@ class EventCleanup(threading.Thread):
Event.thumbnail,
)
.where(
Event.camera.not_in(self.camera_keys),
Event.camera.not_in(self.camera_keys), # type: ignore[arg-type,call-arg,misc]
Event.start_time < expire_after,
Event.label == event.label,
Event.retain_indefinitely == False,
@@ -111,7 +111,7 @@ class EventCleanup(threading.Thread):
# update the clips attribute for the db entry
query = Event.select(Event.id).where(
Event.camera.not_in(self.camera_keys),
Event.camera.not_in(self.camera_keys), # type: ignore[arg-type,call-arg,misc]
Event.start_time < expire_after,
Event.label == event.label,
Event.retain_indefinitely == False,
@@ -218,7 +218,7 @@ class EventCleanup(threading.Thread):
Event.camera,
)
.where(
Event.camera.not_in(self.camera_keys),
Event.camera.not_in(self.camera_keys), # type: ignore[arg-type,call-arg,misc]
Event.start_time < expire_after,
Event.retain_indefinitely == False,
)
@@ -249,7 +249,7 @@ class EventCleanup(threading.Thread):
# update the clips attribute for the db entry
query = Event.select(Event.id).where(
Event.camera.not_in(self.camera_keys),
Event.camera.not_in(self.camera_keys), # type: ignore[arg-type,call-arg,misc]
Event.start_time < expire_after,
Event.retain_indefinitely == False,
)
+2 -89
View File
@@ -105,21 +105,8 @@ class GenAIClient:
debug_save: bool,
activity_context_prompt: str,
response_style: str = "default",
frame_captions: list[str] | None = None,
) -> ReviewMetadata | None:
"""Generate a description for the review item activity.
`frame_captions` holds one caption per thumbnail for the annotated
frame mode; each is sent directly before its frame.
"""
if frame_captions and len(frame_captions) != len(thumbnails):
logger.warning(
"Got %d frame captions for %d thumbnails, sending plain frames",
len(frame_captions),
len(thumbnails),
)
frame_captions = None
"""Generate a description for the review item activity."""
context_prompt = build_review_description_prompt(
review_data,
thumbnails,
@@ -127,7 +114,6 @@ class GenAIClient:
preferred_language,
activity_context_prompt,
response_style,
frame_captions,
)
logger.debug(
@@ -143,30 +129,9 @@ class GenAIClient:
) as f:
f.write(context_prompt)
if frame_captions:
# One file per frame, numbered to match the image it precedes
# (0.txt goes with 0.jpg), so the debug folder replays without
# having to re-derive the mapping.
for index, caption in enumerate(frame_captions):
with open(
os.path.join(
CLIPS_DIR,
"genai-requests",
review_data["id"],
f"{index}.txt",
),
"w",
) as f:
f.write(caption)
response_format = build_review_description_response_format(concerns)
response = self._send(
context_prompt,
thumbnails,
response_format,
image_captions=frame_captions,
)
response = self._send(context_prompt, thumbnails, response_format)
if debug_save and response:
with open(
@@ -304,7 +269,6 @@ class GenAIClient:
images: list[bytes],
response_format: dict | None = None,
enable_thinking: bool = False,
image_captions: list[str] | None = None,
) -> str | None:
"""Submit a request to the provider.
@@ -312,10 +276,6 @@ class GenAIClient:
``supports_toggleable_thinking``. Description-style callers leave it
at the default (off) since synthesis tasks don't benefit from
reasoning traces.
``image_captions`` carries one caption per image, to be placed
immediately before its image so the model can tell the frames apart.
Providers build their request order with ``interleave_images``.
"""
return None
@@ -338,11 +298,6 @@ class GenAIClient:
"""Whether the configured model can generate embeddings via embed()."""
return False
@property
def supports_transcription(self) -> bool:
"""Whether the configured model can transcribe audio via transcribe()."""
return False
def list_models(self) -> list[str]:
"""Return the list of model names available from this provider.
@@ -350,21 +305,6 @@ class GenAIClient:
"""
return []
def list_model_capabilities(self) -> dict[str, dict[str, bool]]:
"""Return capability flags for each model the provider serves.
Only providers whose backend advertises capabilities per model can
populate this; llama.cpp reports input modalities for every model it
serves, so one request describes them all. An empty mapping means "no
per-model information available", and callers fall back to this
client's own capability properties, which describe only the configured
model. A model absent from a non-empty mapping means the same thing.
Returns:
Model name (including aliases) to its capability flags
"""
return {}
def get_context_size(self) -> int:
"""Get the context window size for this provider in tokens."""
return 4096
@@ -396,33 +336,6 @@ class GenAIClient:
)
return []
def transcribe(
self,
audio: bytes,
language: str | None = None,
mime_type: str = "audio/wav",
) -> str | None:
"""Transcribe speech audio to text.
Audio is passed as a self-describing blob rather than raw samples so
every provider receives a container it can declare, and WAV framing
lives in one place instead of in each plugin.
Args:
audio: The encoded audio payload (WAV bytes by default)
language: Optional ISO language hint for the provider
mime_type: Media type of ``audio``
Returns:
The transcript, or None when the provider cannot produce one
"""
logger.warning(
"%s does not support transcription. "
"This method should be overridden by the provider implementation.",
self.__class__.__name__,
)
return None
def chat_with_tools(
self,
messages: list[dict[str, Any]],
-12
View File
@@ -110,12 +110,6 @@ class GenAIClientManager:
name = self._role_map.get(GenAIRoleEnum.embeddings)
return self._get_client(name) if name else None
@property
def transcribe_client(self) -> "GenAIClient | None":
"""Client configured for the transcribe role."""
name = self._role_map.get(GenAIRoleEnum.transcribe)
return self._get_client(name) if name else None
def role_info(self) -> dict[str, dict[str, Any]]:
"""Return the model selected for each configured role and its context size.
@@ -150,11 +144,5 @@ class GenAIClientManager:
"roles": [r.value for r in genai_cfg.roles],
"supports_toggleable_thinking": client.supports_toggleable_thinking,
"supports_embeddings": client.supports_embeddings,
"supports_transcription": client.supports_transcription,
# Capabilities of the configured model are above; this maps every
# model the provider serves to its own, so the UI can react to a
# model selected but not yet saved. Empty when the provider
# cannot report capabilities without loading a model.
"model_capabilities": client.list_model_capabilities(),
}
return result
-13
View File
@@ -13,19 +13,6 @@ overrides what is genuinely Azure-specific:
- Context size: Azure does not expose a per-model ``max_model_len`` field
reliably, so we keep the historical 128K default rather than the
model-name heuristic used by OpenAI.
Transcription is inherited too: :class:`openai.AzureOpenAI` exposes the same
``audio.transcriptions.create``. Two Azure-specific caveats apply when using
the ``transcribe`` role:
- ``model`` must be the Azure *deployment* name, not the underlying model name.
- The ``api-version`` parsed from ``base_url`` must be 2024-06-01 or later;
earlier versions have no transcriptions route and the 404 surfaces only as a
generic provider error.
- Because ``model`` is a deployment name, the inherited check that picks
``languages`` over ``language`` for gpt-transcribe cannot fire unless the
deployment happens to be named after the model. Name the deployment
``gpt-transcribe`` to get the right field, or leave the language on ``auto``.
"""
import logging
+3 -65
View File
@@ -13,14 +13,9 @@ from google.genai.types import FunctionCallingConfigMode
from frigate.config import GenAIProviderEnum
from frigate.genai import GenAIClient, register_genai_provider
from frigate.genai.utils import interleave_images
logger = logging.getLogger(__name__)
# Gemini requests carrying inline data are capped at ~20 MB total; stay well
# under it so the request fails as a log line rather than a 400.
GEMINI_MAX_INLINE_BYTES = 15 * 1024 * 1024
def _decode_thought_signature(value: Any) -> bytes | None:
"""Decode a base64-encoded thought_signature carried across conversation turns."""
@@ -123,16 +118,11 @@ class GeminiClient(GenAIClient):
images: list[bytes],
response_format: dict | None = None,
enable_thinking: bool = False,
image_captions: list[str] | None = None,
) -> str | None:
"""Submit a request to Gemini."""
contents: list[Any] = [
part
if isinstance(part, str)
else types.Part.from_bytes(data=part, mime_type="image/jpeg")
for part in interleave_images(prompt, images, image_captions)
contents = [prompt] + [
types.Part.from_bytes(data=img, mime_type="image/jpeg") for img in images
]
try:
# Merge runtime_options into generation_config if provided
generation_config_dict: dict[str, Any] = {"candidate_count": 1}
@@ -146,7 +136,7 @@ class GeminiClient(GenAIClient):
response = self.provider.models.generate_content(
model=self.genai_config.model,
contents=contents,
contents=contents, # type: ignore[arg-type]
config=types.GenerateContentConfig(
**generation_config_dict,
),
@@ -167,58 +157,6 @@ class GeminiClient(GenAIClient):
return None
return description
@property
def supports_transcription(self) -> bool:
"""Gemini models accept inline audio parts."""
return True
def transcribe(
self,
audio: bytes,
language: str | None = None,
mime_type: str = "audio/wav",
) -> str | None:
"""Transcribe audio by sending it as an inline part alongside a prompt."""
if len(audio) > GEMINI_MAX_INLINE_BYTES:
logger.warning(
"Audio payload of %d bytes exceeds the Gemini inline limit; skipping transcription",
len(audio),
)
return None
prompt = "Transcribe the speech in this audio verbatim. Respond with the transcript only, and with nothing at all if there is no speech."
if language:
prompt += f" The speech is in language '{language}'."
try:
contents: list[Any] = [
prompt,
types.Part.from_bytes(data=audio, mime_type=mime_type),
]
response = self.provider.models.generate_content(
model=self.genai_config.model,
contents=contents,
config=types.GenerateContentConfig(candidate_count=1),
)
except errors.APIError as e:
logger.warning("Gemini returned an error: %s", str(e))
return None
except Exception as e:
logger.warning("An unexpected error occurred with Gemini: %s", str(e))
return None
try:
if response.text is None:
return None
transcript = response.text.strip()
except (ValueError, AttributeError):
# No transcript was generated
return None
return transcript or None
def list_models(self) -> list[str]:
"""Return available model names from Gemini."""
try:
+18 -183
View File
@@ -14,7 +14,7 @@ from PIL import Image
from frigate.config import GenAIProviderEnum
from frigate.genai import GenAIClient, register_genai_provider
from frigate.genai.utils import interleave_images, parse_tool_calls_from_message
from frigate.genai.utils import parse_tool_calls_from_message
logger = logging.getLogger(__name__)
@@ -333,7 +333,6 @@ class LlamaCppClient(GenAIClient):
images: list[bytes],
response_format: dict | None = None,
enable_thinking: bool = False,
image_captions: list[str] | None = None,
) -> str | None:
"""Submit a request to llama.cpp server."""
if self.provider is None:
@@ -343,17 +342,18 @@ class LlamaCppClient(GenAIClient):
return None
try:
content: list[dict[str, Any]] = []
for part in interleave_images(prompt, images, image_captions):
if isinstance(part, str):
content.append({"type": "text", "text": part})
continue
encoded_image = base64.b64encode(part).decode("utf-8")
content = [
{
"type": "text",
"text": prompt,
}
]
for image in images:
encoded_image = base64.b64encode(image).decode("utf-8")
content.append(
{
"type": "image_url",
"image_url": {
"image_url": { # type: ignore[dict-item]
"url": f"data:image/jpeg;base64,{encoded_image}",
},
}
@@ -408,126 +408,6 @@ class LlamaCppClient(GenAIClient):
"""Whether the loaded model supports audio input."""
return self._supports_audio
@property
def supports_transcription(self) -> bool:
"""Audio-capable models can transcribe through chat completions."""
return self._supports_audio
def transcribe(
self,
audio: bytes,
language: str | None = None,
mime_type: str = "audio/wav",
) -> str | None:
"""Transcribe audio through the OpenAI-compatible transcriptions route.
llama.cpp serves /v1/audio/transcriptions for any audio-capable model,
not only a separately loaded whisper (ggml-org/llama.cpp#21863), so it
covers exactly the models supports_transcription detects. It takes the
language as a native multipart field, which is the only thing dedicated
ASR models honor: they read the chat prompt as contextual biasing, so
asking one there to use a language does nothing.
Falls back to chat completions when the server predates that route.
"""
if self.provider is None:
logger.warning(
"llama.cpp provider has not been initialized, audio will not be transcribed. Check your llama.cpp configuration."
)
return None
if not self._supports_audio:
logger.warning(
"llama.cpp model '%s' does not accept audio input",
self.genai_config.model,
)
return None
try:
data = {"model": self.genai_config.model, "response_format": "json"}
if language:
data["language"] = language
response = self._post(
f"{self.provider}/v1/audio/transcriptions",
files={"file": ("audio.wav", audio, mime_type)},
data=data,
timeout=self.timeout,
)
if response.status_code == 404:
logger.debug(
"llama.cpp server has no /v1/audio/transcriptions route, using chat completions"
)
return self._transcribe_via_chat(audio, language)
response.raise_for_status()
result = response.json()
text = result.get("text") if isinstance(result, dict) else None
return str(text).strip() or None if text else None
except Exception as e:
logger.warning("llama.cpp returned an error: %s", str(e))
return None
def _transcribe_via_chat(self, audio: bytes, language: str | None) -> str | None:
"""Transcribe through /v1/chat/completions, for servers without the
transcriptions route.
The _media_marker / multimodal_data convention is an /embeddings-only
protocol, so no marker-refresh retry is needed here.
"""
prompt = "Transcribe the speech in this audio verbatim. Respond with the transcript only, and with nothing at all if there is no speech."
if language:
prompt += f" The speech is in language '{language}'."
try:
encoded_audio = base64.b64encode(audio).decode("utf-8")
payload: dict[str, Any] = {
"model": self.genai_config.model,
"messages": [
{
"role": "user",
"content": [
{"type": "text", "text": prompt},
{
"type": "input_audio",
"input_audio": {
"data": encoded_audio,
"format": "wav",
},
},
],
},
],
**self.provider_options,
}
response = self._post(
f"{self.provider}/v1/chat/completions",
json=payload,
timeout=self.timeout,
)
response.raise_for_status()
result = response.json()
if (
result is not None
and "choices" in result
and len(result["choices"]) > 0
):
choice = result["choices"][0]
if "message" in choice and choice["message"].get("content"):
return str(choice["message"]["content"].strip()) or None
return None
except Exception as e:
logger.warning("llama.cpp returned an error: %s", str(e))
return None
@property
def supports_tools(self) -> bool:
"""Whether the loaded model supports tool/function calling."""
@@ -537,73 +417,28 @@ class LlamaCppClient(GenAIClient):
def supports_toggleable_thinking(self) -> bool:
return self._supports_reasoning
def _fetch_models_data(self) -> list[dict[str, Any]]:
"""Return the raw /v1/models entries, or an empty list if unreachable."""
def list_models(self) -> list[str]:
"""Return available model IDs from the llama.cpp server."""
base_url = self.provider or (
self.genai_config.base_url.rstrip("/")
if self.genai_config.base_url
else None
)
if base_url is None:
return []
try:
response = self._get(f"{base_url}/v1/models", timeout=10)
response.raise_for_status()
data = response.json().get("data", [])
models = []
for m in response.json().get("data", []):
models.append(m.get("id", "unknown"))
for alias in m.get("aliases", []):
models.append(alias)
return sorted(models)
except Exception as e:
logger.warning("Failed to list llama.cpp models: %s", e)
return []
return data if isinstance(data, list) else []
def list_models(self) -> list[str]:
"""Return available model IDs from the llama.cpp server."""
models: set[str] = set()
# llama-server lists the id among the aliases when --alias is set
for m in self._fetch_models_data():
models.add(m.get("id", "unknown"))
models.update(m.get("aliases", []))
return sorted(models)
def list_model_capabilities(self) -> dict[str, dict[str, bool]]:
"""Report input modalities for every model the server serves.
Since ggml-org/llama.cpp#22952 each /v1/models entry carries
architecture.input_modalities, so a single request describes every
model rather than just the configured one. That is what lets the UI
answer "can the model I just picked transcribe" before the config is
saved and a client for it exists.
Models whose entry predates that field are omitted rather than reported
as incapable, so an older server falls back to the /props probe instead
of silently losing capabilities it actually has.
"""
capabilities: dict[str, dict[str, bool]] = {}
for model in self._fetch_models_data():
architecture = model.get("architecture") or {}
modalities = architecture.get("input_modalities")
if not isinstance(modalities, list) or not modalities:
continue
flags = {
"supports_vision": "image" in modalities,
"supports_transcription": "audio" in modalities,
}
names = [model.get("id"), *(model.get("aliases") or [])]
for name in names:
if isinstance(name, str) and name:
capabilities[name] = flags
return capabilities
def get_context_size(self) -> int:
"""Get the context window size for llama.cpp.
+52 -88
View File
@@ -14,7 +14,7 @@ from ollama import ResponseError
from frigate.config import GenAIProviderEnum
from frigate.genai import GenAIClient, register_genai_provider
from frigate.genai.utils import interleave_images, parse_tool_calls_from_message
from frigate.genai.utils import parse_tool_calls_from_message
logger = logging.getLogger(__name__)
@@ -50,28 +50,6 @@ def _extract_ollama_stats(response: Any) -> dict[str, Any] | None:
return stats or None
# Ollama replaces each occurrence of this marker in a message, in order, with
# the next image from the message's images list. Without markers it puts every
# image before the text.
IMAGE_PLACEHOLDER = "[img]"
def _flatten_parts(parts: list[str | bytes]) -> tuple[str, list[bytes] | None]:
"""Collapse ordered text and image parts into Ollama's (content, images)
shape, marking where each image goes so the order survives."""
text: list[str] = []
images: list[bytes] = []
for part in parts:
if isinstance(part, bytes):
text.append(IMAGE_PLACEHOLDER)
images.append(part)
elif part:
text.append(part)
return "\n".join(text), (images or None)
def _normalize_multimodal_content(
content: Any,
) -> tuple[str | None, list[bytes] | None]:
@@ -80,13 +58,13 @@ def _normalize_multimodal_content(
The chat API constructs user messages with content as a list of
``{"type": "text"}`` and ``{"type": "image_url"}`` parts when a tool
returns a live frame. Ollama's SDK requires content to be a string and
images to be passed in a separate field, so images are pulled out and
their positions marked with placeholders.
images to be passed in a separate field, so we extract each.
"""
if not isinstance(content, list):
return content, None
parts: list[str | bytes] = []
text_parts: list[str] = []
images: list[bytes] = []
for part in content:
if not isinstance(part, dict):
continue
@@ -94,20 +72,17 @@ def _normalize_multimodal_content(
if part_type == "text":
text = part.get("text")
if text:
parts.append(str(text))
text_parts.append(str(text))
elif part_type == "image_url":
url = (part.get("image_url") or {}).get("url", "")
if isinstance(url, str) and url.startswith("data:"):
try:
encoded = url.split(",", 1)[1]
parts.append(base64.b64decode(encoded, validate=True))
images.append(base64.b64decode(encoded, validate=True))
except (ValueError, IndexError, binascii.Error) as e:
logger.debug("Failed to decode multimodal image url: %s", e)
if not parts:
return None, None
return _flatten_parts(parts)
return ("\n".join(text_parts) if text_parts else None), (images or None)
@register_genai_provider(GenAIProviderEnum.ollama)
@@ -221,46 +196,58 @@ class OllamaClient(GenAIClient):
images: list[bytes],
response_format: dict | None = None,
enable_thinking: bool = False,
image_captions: list[str] | None = None,
) -> str | None:
"""Submit a request to Ollama through the chat API, the same path the
tool-calling chat uses, with image placeholders keeping any captions
next to their frames."""
"""Submit a request to Ollama"""
if self.provider is None:
logger.warning(
"Ollama provider has not been initialized, a description will not be generated. Check your Ollama configuration."
)
return None
content, message_images = _flatten_parts(
interleave_images(prompt, images, image_captions)
)
message: dict[str, Any] = {"role": "user", "content": content}
if message_images:
message["images"] = message_images
request_params = self._build_request_params(
[message], None, None, enable_thinking=enable_thinking
)
if response_format and response_format.get("type") == "json_schema":
schema = response_format.get("json_schema", {}).get("schema")
if schema:
request_params["format"] = self._clean_schema_for_ollama(schema)
logger.debug(
"Ollama chat request: model=%s, prompt_len=%s, image_count=%s, "
"has_format=%s, think=%s",
self.genai_config.model,
len(prompt),
len(images),
"format" in request_params,
request_params.get("think"),
)
try:
response = self.provider.chat(**request_params)
ollama_options = {
**self.provider_options,
**self.genai_config.runtime_options,
}
if response_format and response_format.get("type") == "json_schema":
schema = response_format.get("json_schema", {}).get("schema")
if schema:
ollama_options["format"] = self._clean_schema_for_ollama(schema)
if self.supports_toggleable_thinking:
ollama_options["think"] = enable_thinking
logger.debug(
"Ollama generate request: model=%s, prompt_len=%s, image_count=%s, "
"has_format=%s, options=%s",
self.genai_config.model,
len(prompt),
len(images) if images else 0,
"format" in ollama_options,
{k: v for k, v in ollama_options.items() if k != "format"},
)
result = self.provider.generate(
self.genai_config.model,
prompt,
images=images if images else None,
**ollama_options,
)
logger.debug(
"Ollama generate response: done=%s, done_reason=%s, eval_count=%s, "
"prompt_eval_count=%s, response_len=%s",
result.get("done"),
result.get("done_reason"),
result.get("eval_count"),
result.get("prompt_eval_count"),
len(result.get("response", "") or ""),
)
response_text = str(result["response"]).strip()
if not response_text:
logger.warning(
"Ollama returned a blank response for model %s (done_reason=%s, "
"eval_count=%s). Check model output, ensure thinking is disabled.",
self.genai_config.model,
result.get("done_reason"),
result.get("eval_count"),
)
return response_text
except (
TimeoutException,
ResponseError,
@@ -270,27 +257,6 @@ class OllamaClient(GenAIClient):
logger.warning("Ollama returned an error: %s", str(e))
return None
logger.debug(
"Ollama chat response: done=%s, done_reason=%s, eval_count=%s, "
"prompt_eval_count=%s",
response.get("done"),
response.get("done_reason"),
response.get("eval_count"),
response.get("prompt_eval_count"),
)
response_text = self._message_from_response(response)["content"] or ""
if not response_text:
logger.warning(
"Ollama returned a blank response for model %s (done_reason=%s, "
"eval_count=%s). Check model output, ensure thinking is disabled.",
self.genai_config.model,
response.get("done_reason"),
response.get("eval_count"),
)
return response_text
def list_models(self) -> list[str]:
"""Return available model names from the Ollama server."""
client = self.provider
@@ -340,8 +306,6 @@ class OllamaClient(GenAIClient):
}
if images:
msg_dict["images"] = images
elif msg.get("images"):
msg_dict["images"] = msg["images"]
if msg.get("tool_call_id"):
msg_dict["tool_call_id"] = msg["tool_call_id"]
if msg.get("name"):
+9 -61
View File
@@ -11,16 +11,9 @@ from openai import OpenAI
from frigate.config import GenAIProviderEnum
from frigate.genai import GenAIClient, register_genai_provider
from frigate.genai.utils import interleave_images
logger = logging.getLogger(__name__)
# gpt-transcribe replaced the singular `language` field with a `languages` array
# and rejects a request that sends both. Older transcription models
# (gpt-4o-transcribe, gpt-4o-mini-transcribe, whisper-1) still take the singular
# form. https://developers.openai.com/api/docs/guides/speech-to-text
_LANGUAGES_ARRAY_MODEL_PREFIX = "gpt-transcribe"
def _stats_from_openai_usage(usage: Any) -> dict[str, Any] | None:
"""Build a stats dict from an OpenAI-compatible usage object."""
@@ -70,21 +63,21 @@ class OpenAIClient(GenAIClient):
images: list[bytes],
response_format: dict | None = None,
enable_thinking: bool = False,
image_captions: list[str] | None = None,
) -> str | None:
"""Submit a request to OpenAI."""
messages_content: list[dict] = []
for part in interleave_images(prompt, images, image_captions):
if isinstance(part, str):
messages_content.append({"type": "text", "text": part})
continue
encoded = base64.b64encode(part).decode("utf-8")
encoded_images = [base64.b64encode(image).decode("utf-8") for image in images]
messages_content: list[dict] = [
{
"type": "text",
"text": prompt,
}
]
for image in encoded_images:
messages_content.append(
{
"type": "image_url",
"image_url": {
"url": f"data:image/jpeg;base64,{encoded}",
"url": f"data:image/jpeg;base64,{image}",
"detail": "low",
},
}
@@ -140,51 +133,6 @@ class OpenAIClient(GenAIClient):
logger.warning("OpenAI returned an error: %s", str(e))
return None
@property
def supports_transcription(self) -> bool:
"""OpenAI exposes /v1/audio/transcriptions for its speech models."""
return True
def transcribe(
self,
audio: bytes,
language: str | None = None,
mime_type: str = "audio/wav",
) -> str | None:
"""Transcribe audio via the OpenAI audio transcriptions endpoint."""
try:
# runtime_options are chat-completion parameters; the transcriptions
# endpoint rejects unknown fields, so they are deliberately not splatted
# in here the way _send() does.
request_params: dict[str, Any] = {
"model": self.genai_config.model,
"file": ("audio.wav", audio, mime_type),
"response_format": "text",
"timeout": self.timeout,
}
if language:
if (
self.genai_config.model.strip()
.lower()
.startswith(_LANGUAGES_ARRAY_MODEL_PREFIX)
):
# not a typed parameter on the SDK method, so it has to ride
# along in extra_body
request_params["extra_body"] = {"languages": [language]}
else:
request_params["language"] = language
result = self.provider.audio.transcriptions.create(**request_params)
except (TimeoutException, Exception) as e:
logger.warning("OpenAI returned an error: %s", str(e))
return None
# response_format="text" yields a bare string, but some compatible
# servers still return the object form
text = result if isinstance(result, str) else getattr(result, "text", None)
return text.strip() if text else None
def list_models(self) -> list[str]:
"""Return available model IDs from the OpenAI-compatible API."""
try:
+2 -15
View File
@@ -59,13 +59,6 @@ def get_review_field_guidelines(response_style: str = "default") -> dict[str, st
}
# Explains the per-frame labels and tracker notes used by the annotated frame
# mode. Neither the notes nor this guidance say whether repeated detections are
# the same subject, since the tracking data cannot tell.
FRAME_ANNOTATION_GUIDANCE = """- Each image below is immediately preceded by a text label giving its frame number and how many seconds into the sequence it was captured. Use these labels to track the order of events and the time between them.
- Some images below are preceded by notes from the camera's object tracker recording what changed at that point: an object being first detected, starting to move, reversing direction, stopping, or no longer being detected. These notes come from tracking data rather than from the images, and they are reliable. Use them to establish how many distinct activities occur and in what order, and describe every one of them."""
def build_review_description_prompt(
review_data: dict[str, Any],
thumbnails: list[bytes],
@@ -73,13 +66,8 @@ def build_review_description_prompt(
preferred_language: str | None,
activity_context_prompt: str,
response_style: str = "default",
frame_captions: list[str] | None = None,
) -> str:
"""Build the prompt for review activity description generation.
When `frame_captions` is set, each caption is sent directly before its
image, so the prompt explains that layout.
"""
"""Build the prompt for review activity description generation."""
def get_concern_prompt() -> str:
if concerns:
@@ -105,7 +93,6 @@ def build_review_description_prompt(
return "\n- (No objects detected)"
fields = get_review_field_guidelines(response_style)
frame_guidance = f"\n{FRAME_ANNOTATION_GUIDANCE}" if frame_captions else ""
return f"""
Your task is to analyze a sequence of images taken in chronological order from a security camera.
@@ -143,7 +130,7 @@ Respond with a JSON object matching the provided schema. Field-specific guidance
## Sequence Details
- Camera: {review_data["camera"]}
- Total frames: {len(thumbnails)} (Frame 1 = earliest, Frame {len(thumbnails)} = latest){frame_guidance}
- Total frames: {len(thumbnails)} (Frame 1 = earliest, Frame {len(thumbnails)} = latest)
- Activity started at {review_data["start"]} and lasted {review_data["duration"]} seconds
- Zones involved: {", ".join(review_data["zones"]) if review_data["zones"] else "None"}
-19
View File
@@ -7,25 +7,6 @@ from typing import Any
logger = logging.getLogger(__name__)
def interleave_images(
prompt: str, images: list[bytes], captions: list[str] | None = None
) -> list[str | bytes]:
"""The prompt, then each image preceded by its caption when one is given.
Providers map the text and image parts onto their own request format, so
every provider sends the same order.
"""
parts: list[str | bytes] = [prompt]
for index, image in enumerate(images):
if captions and index < len(captions):
parts.append(captions[index])
parts.append(image)
return parts
def parse_tool_calls_from_message(
message: dict[str, Any],
) -> list[dict[str, Any]] | None:
+2 -2
View File
@@ -216,11 +216,11 @@ class ExportDebugReplaySource(DebugReplaySource):
"""
def __init__(self, export: Export, duration: float) -> None:
self._camera = export.camera
self._camera = cast(str, export.camera)
# Export.date is declared DateTimeField but Frigate writes raw unix
# timestamps to the column.
self._start_ts = float(cast(Any, export.date))
self._video_path = export.video_path
self._video_path = cast(str, export.video_path)
self._duration = duration
@property
-8
View File
@@ -144,14 +144,6 @@ class LogPipe(threading.Thread):
self.pipeReader.close()
def dump(self) -> None:
if not self.deque:
return
self.logger.log(
self.level,
"The following ffmpeg logs include the last 100 lines prior to exit.",
)
while len(self.deque) > 0:
self.logger.log(self.level, self.deque.popleft())
-2
View File
@@ -109,8 +109,6 @@ class Export(Model):
backref="exports",
column_name="export_case_id",
)
# peewee adds this accessor for the export_case column at runtime
export_case_id: str | None
class ReviewSegment(Model):
+1 -5
View File
@@ -188,12 +188,8 @@ class ImprovedMotionDetector(MotionDetector):
self.config.skip_motion_threshold is not None
and pct_motion > self.config.skip_motion_threshold
):
# recalibrate so we transition to the new background. the frame
# still has to be blended in here, otherwise the background stays
# frozen and every subsequent frame skips as well
# force a recalibration so we transition to the new background
self.calibrating = True
cv2.accumulateWeighted(resized_frame, self.avg_frame, 0.2)
self.motion_frame_count = 0
return []
# once the motion is less than 5% and the number of contours is < 4, assume its calibrated
+3 -3
View File
@@ -186,7 +186,7 @@ class NoticeRegistry:
deleted = (
Notice.delete()
.where(
Notice.kind.in_(camera_kinds),
Notice.kind.in_(camera_kinds), # type: ignore[call-arg, arg-type, misc]
Notice.scope == camera,
)
.execute()
@@ -257,7 +257,7 @@ class NoticeRegistry:
"""Dismissed config and stream check rows, newest first."""
rows = (
Notice.select()
.where(Notice.kind.in_(list(CHECK_KINDS)))
.where(Notice.kind.in_(list(CHECK_KINDS))) # type: ignore[call-arg, arg-type, misc]
.order_by(Notice.dismissed_at.desc())
)
return [{"id": row.id, "dismissed_at": row.dismissed_at} for row in rows]
@@ -351,7 +351,7 @@ class NoticeRegistry:
)
Notice.delete().where(
Notice.kind == kind,
Notice.id.not_in(newest),
Notice.id.not_in(newest), # type: ignore[call-arg, misc]
).execute()
def _bump_occurrences(self, kind: str, count: int, now: float) -> None:
+2 -8
View File
@@ -66,12 +66,6 @@ def get_cache_image_name(camera: str, frame_time: float) -> str:
)
def is_camera_preview_frame(file_name: str, camera: str) -> bool:
"""Check whether a cached preview frame file belongs to the camera."""
# camera names may contain "-", so a prefix match would include "front-door"
return file_name.rsplit("-", 1)[0] == f"preview_{camera}"
def get_most_recent_preview_frame(
camera: str, before: float | None = None
) -> str | None:
@@ -85,7 +79,7 @@ def get_most_recent_preview_frame(
preview_files = [
f
for f in os.listdir(PREVIEW_CACHE_DIR)
if is_camera_preview_frame(f, camera)
if f.startswith(f"preview_{camera}-")
and f.endswith(f".{PREVIEW_FRAME_TYPE}")
]
@@ -282,7 +276,7 @@ class PreviewRecorder:
start_file = f"{file_start}{start_ts}.webp"
for file in sorted(os.listdir(os.path.join(CACHE_DIR, FOLDER_PREVIEW_FRAMES))):
if not is_camera_preview_frame(file, self.camera_name):
if not file.startswith(file_start):
continue
if file < start_file:
+2 -46
View File
@@ -11,50 +11,11 @@ import requests
from numpy import ndarray
from requests.models import Response
from frigate.const import MODEL_CACHE_DIR, PLUS_API_HOST, PLUS_ENV_VAR
from frigate.const import PLUS_API_HOST, PLUS_ENV_VAR
logger = logging.getLogger(__name__)
def add_hailo_alias(model_info: dict[str, Any]) -> dict[str, Any]:
"""Name the hailo detector by its current key as well as its old one.
Frigate+ reports every Hailo model as supporting hailo8l, which this
detector was called before it was renamed to cover every Hailo device.
The old key is kept so an older Frigate still matches the model.
Args:
model_info: A Frigate+ model's metadata, edited in place
Returns:
The same metadata
"""
supported = model_info.get("supportedDetectors")
if supported and "hailo8l" in supported and "hailo" not in supported:
supported.append("hailo")
return model_info
def load_plus_model_info(model_id: str) -> dict[str, Any] | None:
"""Read a Frigate+ model's cached info file.
Args:
model_id: The Frigate+ model id
Returns:
The model info, or None when it has not been cached or cannot be read
"""
try:
with open(os.path.join(MODEL_CACHE_DIR, f"{model_id}.json")) as f:
model_info: dict[str, Any] = json.load(f)
except (OSError, ValueError):
return None
return add_hailo_alias(model_info)
def get_jpg_bytes(image: ndarray, max_dim: int, quality: int) -> bytes:
if image.shape[1] >= image.shape[0]:
width = min(max_dim, image.shape[1])
@@ -280,9 +241,4 @@ class PlusApi:
if not r.ok:
raise Exception(r.text)
models = r.json()
for model in models.get("list") or []:
add_hailo_alias(model)
return models
return r.json()
-22
View File
@@ -1120,18 +1120,6 @@ class OnvifController:
f"Camera {camera_name} is still in ONVIF 'MOVING' status."
)
async def _shutdown(self) -> None:
"""Close the camera sessions and cancel the tasks running on the loop."""
for cam_name in list(self.cams):
await self._close_camera(cam_name)
tasks = [t for t in asyncio.all_tasks() if t is not asyncio.current_task()]
for task in tasks:
task.cancel()
await asyncio.gather(*tasks, return_exceptions=True)
def close(self) -> None:
"""Gracefully shut down the ONVIF controller."""
if not hasattr(self, "loop") or self.loop.is_closed():
@@ -1139,16 +1127,6 @@ class OnvifController:
return
logger.info("Exiting ONVIF controller...")
# anything left open here is garbage collected during interpreter
# shutdown, where its warnings can no longer be logged cleanly
try:
asyncio.run_coroutine_threadsafe(self._shutdown(), self.loop).result(
timeout=5
)
except TimeoutError:
logger.debug("Timed out closing ONVIF sessions")
self.config_subscriber.stop()
def stop_and_cleanup():
+5 -6
View File
@@ -5,7 +5,6 @@ import itertools
import logging
import os
import threading
from collections.abc import Iterable
from multiprocessing.synchronize import Event as MpEvent
from pathlib import Path
from typing import Any
@@ -29,7 +28,7 @@ logger = logging.getLogger(__name__)
def _filter_reviews_for_pass(
reviews: Iterable[Any],
reviews: list[Any],
now: datetime.datetime,
alerts_days: float,
detections_days: float,
@@ -122,14 +121,14 @@ class RecordingCleanup(threading.Thread):
)
maybe_empty_dirs = set()
thumbs_to_delete = list(map(lambda x: x.thumb_path, expired_reviews))
for thumb in thumbs_to_delete:
thumb_path = Path(thumb)
thumbs_to_delete = list(map(lambda x: x[1], expired_reviews))
for thumb_path in thumbs_to_delete:
thumb_path = Path(thumb_path)
thumb_path.unlink(missing_ok=True)
maybe_empty_dirs.add(thumb_path.parent)
max_deletes = 100000
deleted_reviews_list = list(map(lambda x: x.id, expired_reviews))
deleted_reviews_list = list(map(lambda x: x[0], expired_reviews))
for i in range(0, len(deleted_reviews_list), max_deletes):
ReviewSegment.delete().where(
ReviewSegment.id << deleted_reviews_list[i : i + max_deletes]
+5 -10
View File
@@ -14,7 +14,7 @@ from collections.abc import Callable
from dataclasses import dataclass
from enum import Enum
from pathlib import Path
from typing import Any, cast
from typing import Any
import pytz # type: ignore[import-untyped]
from pathvalidate import sanitize_filename
@@ -36,7 +36,6 @@ from frigate.ffmpeg_presets import (
parse_preset_hardware_acceleration_encode,
)
from frigate.models import Export, Previews, Recordings, ReviewSegment
from frigate.output.preview import is_camera_preview_frame
from frigate.util.ffmpeg import run_ffmpeg_with_progress
from frigate.util.ownership import chown_to_runtime
from frigate.util.recording_coverage import (
@@ -51,9 +50,8 @@ from frigate.util.time import is_current_hour
logger = logging.getLogger(__name__)
DEFAULT_TIME_LAPSE_FFMPEG_INPUT_ARGS = "-an"
DEFAULT_TIME_LAPSE_FFMPEG_ARGS = "-vf setpts=0.04*PTS -r 30"
TIMELAPSE_DATA_INPUT_ARGS = "-skip_frame nokey"
TIMELAPSE_DATA_INPUT_ARGS = "-an -skip_frame nokey"
# nginx-vod repackages each stream into fMP4 with a timescale derived from
# that stream's frame rate, and the concat demuxer rescales every input to
@@ -1056,10 +1054,7 @@ class RecordingExporter(threading.Thread):
except DoesNotExist:
return ""
# start_time is a DateTimeField holding a unix timestamp
diff = max(
0.0, float(self.start_time) - float(cast(Any, preview.start_time))
)
diff = max(0.0, float(self.start_time) - float(preview.start_time))
ffmpeg_cmd = [
"/usr/lib/ffmpeg/8.0/bin/ffmpeg", # hardcode path for exports thumbnail due to missing libwebp support
"-hide_banner",
@@ -1099,7 +1094,7 @@ class RecordingExporter(threading.Thread):
fallback_preview = None
for file in sorted(os.listdir(preview_dir)):
if not is_camera_preview_frame(file, self.camera):
if not file.startswith(file_start):
continue
if file < start_file:
@@ -1198,7 +1193,7 @@ class RecordingExporter(threading.Thread):
parse_preset_hardware_acceleration_encode(
self.config.ffmpeg.ffmpeg_path,
hwaccel_args,
f"{self.ffmpeg_input_args} {ffmpeg_input}".strip(),
f"{self.ffmpeg_input_args} -an {ffmpeg_input}".strip(),
f"{self.ffmpeg_output_args} -movflags +faststart".strip(),
EncodeTypeEnum.timelapse,
)
+1 -1
View File
@@ -467,7 +467,7 @@ def get_hardware_temperatures(detector_type: str) -> list[float | None]:
read_temperature(os.path.join(base, apex, "temp"))
for apex in sorted(os.listdir(base))
]
elif detector_type == "hailo":
elif detector_type == "hailo8l":
hailo_temps = get_hailo_temps()
return [hailo_temps[name] for name in sorted(hailo_temps.keys())]
+2 -5
View File
@@ -3,10 +3,8 @@
import logging
import shutil
import threading
from collections.abc import Iterable
from multiprocessing.synchronize import Event as MpEvent
from pathlib import Path
from typing import Any, cast
from peewee import SQL, Case, fn
@@ -193,15 +191,14 @@ class StorageMaintainer(threading.Thread):
stream_usages = {
row["stream_type"]: row["usage"] or 0
for row in cast(
Iterable[dict[str, Any]],
for row in (
Recordings.select(
Recordings.stream_type,
fn.SUM(Recordings.segment_size).alias("usage"),
)
.where(Recordings.camera == camera, Recordings.segment_size != 0)
.group_by(Recordings.stream_type)
.dicts(),
.dicts()
)
}
stream_bandwidths = self.camera_storage_stats.get(camera, {}).get(
@@ -1,81 +0,0 @@
"""Tests that /auth only accepts a JWT whose role is still in the config."""
import os
import time
from unittest.mock import MagicMock, Mock, patch
from frigate.api.auth import create_encoded_jwt
from frigate.api.fastapi_app import create_fastapi_app
from frigate.config import FrigateConfig
from frigate.config.camera.updater import CameraConfigUpdatePublisher
from frigate.const import JWT_SECRET_ENV_VAR
from frigate.models import Event, Recordings, ReviewSegment
from frigate.test.http_api.base_http_test import AuthTestClient, BaseTestHttp
@patch.dict(os.environ, {JWT_SECRET_ENV_VAR: "test-secret"})
class TestAuthJwtRole(BaseTestHttp):
def setUp(self):
super().setUp(models=[Event, Recordings, ReviewSegment])
self.minimal_config = {
"mqtt": {"host": "mqtt"},
"auth": {"enabled": True, "roles": {"garage": ["front_door"]}},
"networking": {"listen": {"internal": 5000, "external": 8971}},
"cameras": {
"front_door": {
"ffmpeg": {
"inputs": [
{"path": "rtsp://10.0.0.1:554/video", "roles": ["detect"]}
]
},
"detect": {
"height": 1080,
"width": 1920,
"fps": 5,
},
}
},
}
def _create_app(self):
mock_publisher = Mock(spec=CameraConfigUpdatePublisher)
mock_publisher.publisher = MagicMock()
return create_fastapi_app(
FrigateConfig(**self.minimal_config),
self.db,
None,
None,
None,
None,
None,
None,
mock_publisher,
None,
enforce_default_admin=False,
)
def _auth(self, app, role: str):
token = create_encoded_jwt("bob", role, int(time.time()) + 3600, app.jwt_token)
with AuthTestClient(app) as client:
return client.get(
"/auth",
headers={
"x-server-port": "8971",
"authorization": f"Bearer {token}",
},
)
def test_configured_role_is_accepted(self):
resp = self._auth(self._create_app(), "garage")
self.assertEqual(resp.status_code, 202)
self.assertEqual(resp.headers["remote-user"], "bob")
self.assertEqual(resp.headers["remote-role"], "garage")
def test_role_removed_from_config_is_rejected(self):
resp = self._auth(self._create_app(), "removed_role")
self.assertEqual(resp.status_code, 401)
self.assertNotIn("remote-role", resp.headers)
@@ -1,193 +0,0 @@
"""Tests for renaming and deleting a zone through config_set's JSON body."""
import os
import tempfile
from unittest.mock import MagicMock, Mock, patch
import ruamel.yaml
from fastapi import Request
from frigate.api.auth import get_allowed_cameras_for_filter, get_current_user
from frigate.api.fastapi_app import create_fastapi_app
from frigate.config import FrigateConfig
from frigate.config.camera.updater import CameraConfigUpdatePublisher
from frigate.models import Event, Recordings, ReviewSegment
from frigate.test.http_api.base_http_test import AuthTestClient, BaseTestHttp
class TestConfigSetZones(BaseTestHttp):
def setUp(self):
super().setUp(models=[Event, Recordings, ReviewSegment])
self.minimal_config = {
"mqtt": {"host": "mqtt"},
"profiles": {"armed": {"friendly_name": "Armed"}},
"snapshots": {"required_zones": ["driveway"]},
"cameras": {
"front": {
"ffmpeg": {
"inputs": [
{"path": "rtsp://10.0.0.1:554/video", "roles": ["detect"]}
]
},
"detect": {"height": 1080, "width": 1920, "fps": 5},
"zones": {
"driveway": {
"coordinates": "0,0,1,0,1,1",
"inertia": 5,
"filters": {"person": {"min_area": 5000}},
},
"porch": {"coordinates": "0,0,0.5,0,0.5,0.5"},
},
"review": {
"alerts": {
"labels": ["person"],
"required_zones": ["driveway"],
},
},
"mqtt": {"required_zones": ["driveway", "porch"]},
"profiles": {
"armed": {
"zones": {
"driveway": {
"coordinates": "0,0,1,0,1,1",
"objects": ["car"],
},
},
},
},
},
},
}
yaml = ruamel.yaml.YAML()
with tempfile.NamedTemporaryFile(
mode="w", suffix=".yml", delete=False
) as config_file:
yaml.dump(self.minimal_config, config_file)
self.config_path = config_file.name
self.addCleanup(os.unlink, self.config_path)
def _create_app(self):
publisher = Mock(spec=CameraConfigUpdatePublisher)
publisher.publisher = MagicMock()
app = create_fastapi_app(
FrigateConfig(**self.minimal_config),
self.db,
None,
None,
None,
None,
None,
None,
publisher,
None,
enforce_default_admin=False,
)
async def mock_get_current_user(request: Request):
return {
"username": request.headers.get("remote-user"),
"role": request.headers.get("remote-role"),
}
async def mock_get_allowed_cameras_for_filter(request: Request):
return ["front"]
app.dependency_overrides[get_current_user] = mock_get_current_user
app.dependency_overrides[get_allowed_cameras_for_filter] = (
mock_get_allowed_cameras_for_filter
)
return app
def _put(self, camera_data: dict):
with patch("frigate.api.app.find_config_file", return_value=self.config_path):
with AuthTestClient(self._create_app()) as client:
return client.put(
"/config/set",
json={
"config_data": {"cameras": {"front": camera_data}},
"requires_restart": 0,
"update_topic": "config/cameras/front/zones",
},
)
def _front(self) -> dict:
with open(self.config_path) as f:
return ruamel.yaml.YAML().load(f)["cameras"]["front"]
def test_rename_moves_zone_references_in_one_request(self):
"""The new base zone exists when validation checks the moved override."""
resp = self._put(
{
"zones": {
"driveway": None,
"front_drive": {
"coordinates": "0,0,1,0,1,1",
"enabled": True,
# as /api/config returns them, with defaults filled in
"filters": {
"person": {
"min_area": 5000,
"max_area": 24000000,
"min_ratio": 0.0,
"max_ratio": 24000000.0,
"threshold": 0.7,
"min_score": 0.5,
"mask": {},
}
},
"inertia": 5,
},
},
"review": {"alerts": {"required_zones": ["front_drive"]}},
# inherited from the global snapshots config, so no camera key
"snapshots": {"required_zones": ["front_drive"]},
"mqtt": {"required_zones": ["front_drive", "porch"]},
"profiles": {
"armed": {
"zones": {
"driveway": None,
"front_drive": {
"coordinates": "0,0,1,0,1,1",
"objects": ["car"],
},
},
},
},
}
)
self.assertEqual(resp.status_code, 200, resp.json())
front = self._front()
self.assertEqual(list(front["zones"]), ["porch", "front_drive"])
self.assertEqual(front["zones"]["front_drive"]["inertia"], 5)
self.assertEqual(
front["zones"]["front_drive"]["filters"]["person"]["min_area"], 5000
)
self.assertEqual(front["review"]["alerts"]["labels"], ["person"])
self.assertEqual(front["review"]["alerts"]["required_zones"], ["front_drive"])
self.assertEqual(front["snapshots"]["required_zones"], ["front_drive"])
self.assertEqual(front["mqtt"]["required_zones"], ["front_drive", "porch"])
self.assertEqual(
dict(front["profiles"]["armed"]["zones"]),
{"front_drive": {"coordinates": "0,0,1,0,1,1", "objects": ["car"]}},
)
def test_delete_empties_lists_without_dropping_their_section(self):
resp = self._put(
{
"zones": {"driveway": None},
"review": {"alerts": {"required_zones": []}},
"snapshots": {"required_zones": []},
"mqtt": {"required_zones": ["porch"]},
"profiles": {"armed": {"zones": {"driveway": None}}},
}
)
self.assertEqual(resp.status_code, 200, resp.json())
front = self._front()
self.assertEqual(list(front["zones"]), ["porch"])
self.assertEqual(front["review"]["alerts"]["labels"], ["person"])
self.assertEqual(front["review"]["alerts"]["required_zones"], [])
self.assertEqual(front["mqtt"]["required_zones"], ["porch"])
self.assertNotIn("driveway", front["profiles"]["armed"]["zones"] or {})
-90
View File
@@ -1,90 +0,0 @@
"""Tests for audio maintainer spawning on runtime camera adds."""
import threading
import unittest
from unittest.mock import MagicMock, patch
from frigate.config import FrigateConfig
from frigate.events.audio import AudioProcessor
def _build_config() -> FrigateConfig:
return FrigateConfig(
**{
"mqtt": {"host": "mqtt"},
"cameras": {
"front_door": {
"ffmpeg": {
"inputs": [
{
"path": "rtsp://10.0.0.1:554/video",
"roles": ["detect", "audio"],
}
]
},
"audio": {"enabled": True},
}
},
}
)
class TestAudioSpawnIfNeeded(unittest.TestCase):
def _processor(self, config: FrigateConfig, camera_metrics: dict) -> AudioProcessor:
processor = AudioProcessor.__new__(AudioProcessor)
processor.config = config
processor.camera_metrics = camera_metrics
processor.embeddings_metrics = MagicMock()
processor.audio_threads = {}
processor.transcription_model_runner = None
processor.genai_manager = None
processor.stop_event = threading.Event()
processor.logger = MagicMock()
return processor
def test_skips_camera_whose_metrics_have_not_been_created_yet(self):
"""The camera maintainer creates metrics on its own poll of the add
update, so the audio processor can see the camera first."""
config = _build_config()
processor = self._processor(config, {})
with patch("frigate.events.audio.AudioEventMaintainer") as maintainer:
processor.spawn_if_needed(config.cameras["front_door"])
maintainer.assert_not_called()
assert processor.audio_threads == {}
def test_spawns_once_metrics_exist_and_passes_the_metrics_object(self):
config = _build_config()
metrics = object()
processor = self._processor(config, {"front_door": metrics})
with patch("frigate.events.audio.AudioEventMaintainer") as maintainer:
processor.spawn_if_needed(config.cameras["front_door"])
assert maintainer.call_args.args[2] is metrics
assert "front_door" in processor.audio_threads
maintainer.return_value.start.assert_called_once()
def test_does_not_respawn_for_a_camera_already_running(self):
config = _build_config()
processor = self._processor(config, {"front_door": object()})
processor.audio_threads["front_door"] = MagicMock()
with patch("frigate.events.audio.AudioEventMaintainer") as maintainer:
processor.spawn_if_needed(config.cameras["front_door"])
maintainer.assert_not_called()
def test_skips_camera_whose_ffmpeg_update_has_not_arrived_yet(self):
"""The add update carries the camera before its ffmpeg inputs are
applied, so the audio role can be briefly missing."""
config = _build_config()
camera = config.cameras["front_door"]
camera.ffmpeg.inputs[0].roles = ["detect"]
processor = self._processor(config, {"front_door": object()})
with patch("frigate.events.audio.AudioEventMaintainer") as maintainer:
processor.spawn_if_needed(camera)
maintainer.assert_not_called()
@@ -1,286 +0,0 @@
"""Config validation and the historical path for the audio_transcription GenAI backend."""
import unittest
from copy import deepcopy
from unittest.mock import MagicMock, patch
from pydantic import ValidationError
from frigate.config import FrigateConfig
from frigate.config.camera.genai import GenAIConfig, GenAIRoleEnum
from frigate.config.classification import AudioTranscriptionModelEnum
from frigate.const import UPDATE_EVENT_DESCRIPTION
from frigate.data_processing.post.audio_transcription import (
AudioTranscriptionPostProcessor,
)
from frigate.data_processing.types import PostProcessDataEnum
class TestAudioTranscriptionGenAIConfig(unittest.TestCase):
def setUp(self):
self.base = {
"mqtt": {"host": "mqtt"},
"cameras": {
"back": {
"ffmpeg": {
"inputs": [
{
"path": "rtsp://10.0.0.1:554/video",
"roles": ["detect", "audio"],
}
]
},
"detect": {"height": 1080, "width": 1920, "fps": 5},
"audio": {"enabled": True},
}
},
}
def _config(self, **overrides) -> dict:
config = deepcopy(self.base)
config.update(deepcopy(overrides))
return config
def _provider(self, roles: list[str]) -> dict:
return {
"whisper_cloud": {
"provider": "openai",
"model": "gpt-4o-transcribe",
"api_key": "k",
"roles": roles,
}
}
def test_default_model_is_whisper_enum(self):
config = FrigateConfig(**self._config())
self.assertEqual(
config.audio_transcription.model, AudioTranscriptionModelEnum.whisper
)
def test_whisper_string_coerces_to_enum(self):
config = FrigateConfig(
**self._config(audio_transcription={"enabled": True, "model": "whisper"})
)
self.assertIsInstance(
config.audio_transcription.model, AudioTranscriptionModelEnum
)
def test_provider_name_stays_a_string(self):
config = FrigateConfig(
**self._config(
genai=self._provider(["transcribe"]),
audio_transcription={"enabled": True, "model": "whisper_cloud"},
)
)
self.assertNotIsInstance(
config.audio_transcription.model, AudioTranscriptionModelEnum
)
self.assertEqual(config.audio_transcription.model, "whisper_cloud")
def test_unspecified_model_falls_back_to_whisper(self):
"""An empty value must not read as a GenAI provider that resolves to no client."""
for value in (None, "", " "):
with self.subTest(repr(value)):
config = FrigateConfig(
**self._config(
audio_transcription={"enabled": True, "model": value}
)
)
self.assertIs(
config.audio_transcription.model,
AudioTranscriptionModelEnum.whisper,
)
def test_missing_genai_key_raises(self):
with self.assertRaises(ValidationError) as ctx:
FrigateConfig(
**self._config(
audio_transcription={"enabled": True, "model": "nope"},
)
)
self.assertIn("is not a valid GenAI config key", str(ctx.exception))
def test_provider_without_role_raises(self):
with self.assertRaises(ValidationError) as ctx:
FrigateConfig(
**self._config(
genai=self._provider(["descriptions"]),
audio_transcription={"enabled": True, "model": "whisper_cloud"},
)
)
self.assertIn("must have 'transcribe' in its roles", str(ctx.exception))
def test_global_off_camera_on_still_validates(self):
"""Global-off/camera-on is a supported deployment and must not skip the check."""
config = self._config(
genai=self._provider(["descriptions"]),
audio_transcription={"enabled": False, "model": "whisper_cloud"},
)
config["cameras"]["back"]["audio_transcription"] = {"enabled": True}
with self.assertRaises(ValidationError) as ctx:
FrigateConfig(**config)
self.assertIn("must have 'transcribe' in its roles", str(ctx.exception))
def test_disabled_transcription_skips_validation(self):
config = FrigateConfig(
**self._config(audio_transcription={"enabled": False, "model": "nope"})
)
self.assertEqual(config.audio_transcription.model, "nope")
def test_camera_level_model_is_rejected(self):
config = self._config()
config["cameras"]["back"]["audio_transcription"] = {
"enabled": True,
"model": "whisper",
}
with self.assertRaises(ValidationError):
FrigateConfig(**config)
def test_default_roles_do_not_include_transcribe(self):
"""Backward compatibility: existing providers must not silently claim it."""
genai = GenAIConfig(provider="openai", model="gpt-4o")
self.assertNotIn(GenAIRoleEnum.transcribe, genai.roles)
def test_two_providers_claiming_transcribe_raises(self):
genai = self._provider(["transcribe"])
genai["other"] = {
"provider": "gemini",
"model": "gemini-2.0-flash",
"api_key": "k",
"roles": ["transcribe"],
}
with self.assertRaises(ValidationError) as ctx:
FrigateConfig(
**self._config(
genai=genai,
audio_transcription={"enabled": True, "model": "whisper_cloud"},
)
)
self.assertIn("each role must have", str(ctx.exception))
def test_transcribe_rejected_on_provider_without_audio_input(self):
with self.assertRaises(ValidationError) as ctx:
GenAIConfig(provider="ollama", model="llava", roles=["transcribe"])
self.assertIn("does not support audio input", str(ctx.exception))
class TestAudioTranscriptionPostProcessorGenAI(unittest.TestCase):
"""The recorded-speech path must reach the provider and skip the local model."""
def setUp(self):
self.config = FrigateConfig(
**{
"mqtt": {"host": "mqtt"},
"genai": {
"whisper_cloud": {
"provider": "openai",
"model": "gpt-4o-transcribe",
"api_key": "k",
"roles": ["transcribe"],
}
},
"audio_transcription": {
"enabled": True,
"model": "whisper_cloud",
"language": "en",
},
"cameras": {
"back": {
"ffmpeg": {
"inputs": [
{
"path": "rtsp://10.0.0.1:554/video",
"roles": ["detect", "audio"],
}
]
},
"detect": {"height": 1080, "width": 1920, "fps": 5},
"audio": {"enabled": True},
}
},
}
)
self.client = MagicMock()
self.client.transcribe.return_value = "recorded speech"
self.manager = MagicMock()
self.manager.transcribe_client = self.client
self.requestor = MagicMock()
self.processor = AudioTranscriptionPostProcessor(
self.config,
self.requestor,
MagicMock(),
MagicMock(),
self.manager,
)
def _process(self):
self.processor.process_data(
{
"event_id": "1234.5-abc",
"camera": "back",
"event": {
"id": "1234.5-abc",
"camera": "back",
"start_time": 100.0,
"end_time": 110.0,
"data": {},
},
},
PostProcessDataEnum.tracked_object,
)
def test_local_recognizer_is_never_built(self):
self.assertTrue(self.processor._use_genai)
self.assertIsNone(self.processor.recognizer)
def test_audio_bytes_and_language_reach_the_client(self):
with patch(
"frigate.data_processing.post.audio_transcription.get_audio_from_recording",
return_value=b"RIFF....WAVE",
):
self._process()
self.client.transcribe.assert_called_once()
self.assertEqual(self.client.transcribe.call_args.args[0], b"RIFF....WAVE")
self.assertEqual(self.client.transcribe.call_args.kwargs["language"], "en")
def test_transcript_is_published_as_the_description(self):
with patch(
"frigate.data_processing.post.audio_transcription.get_audio_from_recording",
return_value=b"RIFF....WAVE",
):
self._process()
topics = [call.args[0] for call in self.requestor.send_data.call_args_list]
self.assertIn(UPDATE_EVENT_DESCRIPTION, topics)
payload = next(
call.args[1]
for call in self.requestor.send_data.call_args_list
if call.args[0] == UPDATE_EVENT_DESCRIPTION
)
self.assertEqual(payload["description"], "recorded speech")
self.assertEqual(payload["id"], "1234.5-abc")
def test_missing_client_publishes_nothing(self):
self.manager.transcribe_client = None
with patch(
"frigate.data_processing.post.audio_transcription.get_audio_from_recording",
return_value=b"RIFF....WAVE",
):
self._process()
self.requestor.send_data.assert_not_called()
if __name__ == "__main__":
unittest.main()
@@ -1,288 +0,0 @@
"""Live GenAI transcription: sliding overlapped windows and their lifecycle."""
import io
import threading
import unittest
import wave
from unittest.mock import MagicMock, patch
import numpy as np
from frigate.config import FrigateConfig
from frigate.const import AUDIO_DURATION, AUDIO_SAMPLE_RATE
from frigate.data_processing.real_time.audio_transcription import (
GENAI_WINDOW_CHUNKS,
AudioTranscriptionRealTimeProcessor,
)
CHUNK_SAMPLES = int(round(AUDIO_DURATION * AUDIO_SAMPLE_RATE))
def _chunk(amplitude: int) -> np.ndarray:
"""One audio-detector-sized chunk of int16 samples at a constant amplitude."""
return np.full(CHUNK_SAMPLES, amplitude, dtype=np.int16)
class TestLiveGenAITranscription(unittest.TestCase):
def setUp(self):
self.config = FrigateConfig(
**{
"mqtt": {"host": "mqtt"},
"genai": {
"whisper_cloud": {
"provider": "openai",
"model": "gpt-4o-transcribe",
"api_key": "k",
"roles": ["transcribe"],
}
},
"audio_transcription": {
"enabled": True,
"model": "whisper_cloud",
"language": "en",
},
"cameras": {
"back": {
"ffmpeg": {
"inputs": [
{
"path": "rtsp://10.0.0.1:554/video",
"roles": ["detect", "audio"],
}
]
},
"detect": {"height": 1080, "width": 1920, "fps": 5},
"audio": {"enabled": True},
}
},
}
)
self.client = MagicMock()
self.client.transcribe.return_value = "hello"
self.manager = MagicMock()
self.manager.transcribe_client = self.client
self.requestor = MagicMock()
self.processor = AudioTranscriptionRealTimeProcessor(
config=self.config,
camera_config=self.config.cameras["back"],
requestor=self.requestor,
model_runner=None,
metrics=MagicMock(),
stop_event=threading.Event(),
genai_manager=self.manager,
)
def _feed(self, chunk: np.ndarray):
return (
self.processor._AudioTranscriptionRealTimeProcessor__process_audio_stream(
chunk
)
)
def _sent_wav(self, call_index: int) -> wave.Wave_read:
payload = self.client.transcribe.call_args_list[call_index].args[0]
return wave.open(io.BytesIO(payload), "rb")
def test_uses_genai_path(self):
self.assertTrue(self.processor._use_genai)
def test_first_chunk_does_not_transcribe(self):
self.assertIsNone(self._feed(_chunk(4000)))
self.client.transcribe.assert_not_called()
def test_full_window_transcribes_two_chunks(self):
self._feed(_chunk(4000))
result = self._feed(_chunk(4000))
self.assertEqual(result, ("hello", False))
self.client.transcribe.assert_called_once()
self.assertEqual(
self.client.transcribe.call_args.kwargs["language"],
"en",
)
with self._sent_wav(0) as wav:
self.assertEqual(wav.getnchannels(), 1)
self.assertEqual(wav.getsampwidth(), 2)
self.assertEqual(wav.getframerate(), AUDIO_SAMPLE_RATE)
self.assertEqual(wav.getnframes(), CHUNK_SAMPLES * GENAI_WINDOW_CHUNKS)
def test_window_slides_with_overlap(self):
"""The third chunk's window is chunks 2+3, not 3 alone and not 1+2+3."""
self.client.transcribe.side_effect = ["one two", "two three"]
self._feed(_chunk(1000))
self._feed(_chunk(2000))
result = self._feed(_chunk(3000))
self.assertEqual(self.client.transcribe.call_count, 2)
with self._sent_wav(1) as wav:
self.assertEqual(wav.getnframes(), CHUNK_SAMPLES * GENAI_WINDOW_CHUNKS)
samples = np.frombuffer(wav.readframes(wav.getnframes()), dtype=np.int16)
self.assertEqual(samples[0], 2000)
self.assertEqual(samples[-1], 3000)
# the shared "two" appears once
self.assertEqual(result, ("one two three", False))
def test_silent_window_is_not_uploaded(self):
self._feed(_chunk(0))
self.assertIsNone(self._feed(_chunk(0)))
self.client.transcribe.assert_not_called()
def test_silent_window_ends_a_pending_utterance(self):
self._feed(_chunk(4000))
self._feed(_chunk(4000))
# the gate covers the whole window, so it takes GENAI_WINDOW_CHUNKS
# silent chunks to push the last speech out of it
self._feed(_chunk(0))
self.assertEqual(self._feed(_chunk(0)), ("hello", True))
def test_empty_transcript_ends_a_pending_utterance(self):
self.client.transcribe.side_effect = ["hello", ""]
self._feed(_chunk(4000))
self._feed(_chunk(4000))
self.assertEqual(self._feed(_chunk(4000)), ("hello", True))
def test_empty_transcript_with_nothing_pending_returns_none(self):
self.client.transcribe.return_value = ""
self._feed(_chunk(4000))
self.assertIsNone(self._feed(_chunk(4000)))
def test_reset_clears_committed_text_and_window(self):
self._feed(_chunk(4000))
self._feed(_chunk(4000))
self.processor.reset()
self.assertEqual(self.processor._genai_committed, "")
self.assertEqual(len(self.processor._genai_window), 0)
# a fresh window is required again before the next request
self.client.transcribe.reset_mock()
self._feed(_chunk(4000))
self.client.transcribe.assert_not_called()
def test_check_unload_model_clears_once_then_is_idempotent(self):
self._feed(_chunk(4000))
self._feed(_chunk(4000))
self.processor.check_unload_model()
self.requestor.send_data.assert_called_once_with("back/audio/transcription", "")
self.assertEqual(self.processor._genai_committed, "")
self.assertEqual(len(self.processor._genai_window), 0)
self.processor.check_unload_model()
self.requestor.send_data.assert_called_once()
def test_build_recognizer_never_loads_a_local_model(self):
with patch(
"frigate.data_processing.real_time.audio_transcription.FasterWhisperASR"
) as whisper:
self.processor._AudioTranscriptionRealTimeProcessor__build_recognizer()
whisper.assert_not_called()
self.assertIsNone(self.processor.stream)
def test_clear_audio_recognizer_request_only_resets(self):
self._feed(_chunk(4000))
self._feed(_chunk(4000))
with patch.object(
self.processor,
"_AudioTranscriptionRealTimeProcessor__build_recognizer",
) as build:
result = self.processor.handle_request("clear_audio_recognizer", {})
build.assert_not_called()
self.assertTrue(result["success"])
self.assertEqual(self.processor._genai_committed, "")
def test_missing_client_logs_and_returns_none(self):
self.manager.transcribe_client = None
self.assertIsNone(self._feed(_chunk(4000)))
self.assertIsNone(self._feed(_chunk(4000)))
def test_dropped_audio_discards_the_buffered_window(self):
"""A gap in the stream must not be spliced into a single window.
Dropping a queued chunk leaves the next one non-adjacent to what is
buffered, so concatenating them would hand the provider audio with a
hole in it and break the 50% overlap the stitcher relies on.
"""
self._feed(_chunk(4000))
self.assertEqual(len(self.processor._genai_window), 1)
# the producer discards a chunk while the consumer is blocked
self.processor._audio_dropped.set()
self._feed(_chunk(5000))
# the buffered chunk was discarded, so this one starts a fresh window
self.assertEqual(len(self.processor._genai_window), 1)
self.client.transcribe.assert_not_called()
# and the window that does go out holds only contiguous audio
self._feed(_chunk(5000))
self.client.transcribe.assert_called_once()
with self._sent_wav(0) as wav:
samples = np.frombuffer(wav.readframes(wav.getnframes()), dtype=np.int16)
self.assertEqual(wav.getnframes(), CHUNK_SAMPLES * GENAI_WINDOW_CHUNKS)
self.assertTrue((samples == 5000).all(), "window spliced across the gap")
def test_dropped_audio_ends_a_pending_utterance(self):
"""Committed text cannot be stitched across missing speech."""
self._feed(_chunk(4000))
self._feed(_chunk(4000))
self.assertEqual(self.processor._genai_committed, "hello")
self.processor._audio_dropped.set()
self.assertEqual(self._feed(_chunk(4000)), ("hello", True))
self.assertEqual(len(self.processor._genai_window), 0)
def test_drop_flag_is_consumed_once(self):
self.processor._audio_dropped.set()
self._feed(_chunk(4000))
self.assertFalse(self.processor._audio_dropped.is_set())
def test_full_queue_flags_a_drop(self):
for i in range(self.processor.audio_queue.maxsize + 1):
self.processor.process_audio({"id": "back_audio"}, _chunk(i + 1))
self.assertTrue(self.processor._audio_dropped.is_set())
def test_queue_is_bounded_and_drops_oldest(self):
maxsize = self.processor.audio_queue.maxsize
self.assertGreater(maxsize, 0)
for i in range(maxsize + 5):
self.processor.process_audio({"id": "back_audio"}, _chunk(i + 1))
self.assertEqual(self.processor.audio_queue.qsize(), maxsize)
# the newest chunk survived, the oldest did not
remaining = []
while not self.processor.audio_queue.empty():
remaining.append(self.processor.audio_queue.get_nowait()[1][0])
self.assertEqual(remaining[-1], maxsize + 5)
self.assertNotIn(1, remaining)
if __name__ == "__main__":
unittest.main()
-176
View File
@@ -1,176 +0,0 @@
"""Tests for the WAV helpers and transcript stitcher in frigate.util.audio."""
import io
import struct
import unittest
import wave
import numpy as np
from frigate.const import AUDIO_SAMPLE_RATE
from frigate.util.audio import fix_wav_header, pcm16_to_wav, stitch_transcripts
def _wav(samples: np.ndarray, sample_rate: int = AUDIO_SAMPLE_RATE) -> bytes:
buffer = io.BytesIO()
with wave.open(buffer, "wb") as out:
out.setnchannels(1)
out.setsampwidth(2)
out.setframerate(sample_rate)
out.writeframes(samples.tobytes())
return buffer.getvalue()
class TestPcm16ToWav(unittest.TestCase):
def test_round_trips_through_wave(self):
samples = np.arange(-1000, 1000, dtype=np.int16)
with wave.open(io.BytesIO(pcm16_to_wav(samples)), "rb") as wav:
self.assertEqual(wav.getnchannels(), 1)
self.assertEqual(wav.getsampwidth(), 2)
self.assertEqual(wav.getframerate(), AUDIO_SAMPLE_RATE)
self.assertEqual(wav.getnframes(), samples.size)
decoded = np.frombuffer(wav.readframes(wav.getnframes()), dtype=np.int16)
np.testing.assert_array_equal(decoded, samples)
def test_casts_non_int16_input(self):
samples = np.array([0.0, 100.0, -100.0], dtype=np.float32)
with wave.open(io.BytesIO(pcm16_to_wav(samples)), "rb") as wav:
decoded = np.frombuffer(wav.readframes(wav.getnframes()), dtype=np.int16)
np.testing.assert_array_equal(decoded, np.array([0, 100, -100], np.int16))
def test_honors_sample_rate(self):
with wave.open(
io.BytesIO(pcm16_to_wav(np.zeros(4, np.int16), 8000)), "rb"
) as w:
self.assertEqual(w.getframerate(), 8000)
class TestFixWavHeader(unittest.TestCase):
def test_rewrites_placeholder_sizes(self):
samples = np.arange(64, dtype=np.int16)
data = bytearray(_wav(samples))
# ffmpeg piping to non-seekable stdout leaves both sizes unpatched
struct.pack_into("<I", data, 4, 0xFFFFFFFF)
data_offset = data.index(b"data")
struct.pack_into("<I", data, data_offset + 4, 0xFFFFFFFF)
fixed = fix_wav_header(bytes(data))
self.assertEqual(struct.unpack_from("<I", fixed, 4)[0], len(fixed) - 8)
self.assertEqual(
struct.unpack_from("<I", fixed, data_offset + 4)[0],
len(fixed) - (data_offset + 8),
)
with wave.open(io.BytesIO(fixed), "rb") as wav:
self.assertEqual(wav.getnframes(), samples.size)
def test_leaves_a_well_formed_header_alone(self):
data = _wav(np.arange(32, dtype=np.int16))
self.assertEqual(fix_wav_header(data), data)
def test_non_riff_payload_passes_through(self):
self.assertEqual(fix_wav_header(b"not a wav"), b"not a wav")
self.assertEqual(fix_wav_header(b""), b"")
class TestStitchTranscripts(unittest.TestCase):
def test_table(self):
cases = [
# (committed, incoming, expected, description)
(
"the quick brown",
"brown fox jumps",
"the quick brown fox jumps",
"one word",
),
(
"and then the quick brown",
"the quick brown fox",
"and then the quick brown fox",
"multi word",
),
(
"hello there",
"general kenobi",
"hello there general kenobi",
"no overlap",
),
("the quick brown fox", "brown fox", "the quick brown fox", "contained"),
("", "first words", "first words", "empty committed"),
("already here", "", "already here", "empty incoming"),
(
" spaced out ",
"out again",
"spaced out again",
"whitespace normalized",
),
]
for committed, incoming, expected, description in cases:
with self.subTest(description):
self.assertEqual(stitch_transcripts(committed, incoming), expected)
def test_overlap_found_mid_window(self):
"""The shared run is rarely at the start of the new window.
The provider re-transcribes the overlapping audio independently and
often renders its first word differently, so anchoring the match to the
start of the incoming window duplicates the whole phrase.
"""
self.assertEqual(
stitch_transcripts(
"this is just gonna be a fun time", "It's gonna be a fun time."
),
"this is just gonna be a fun time",
)
def test_overlap_longer_than_five_words(self):
"""The cap is bounded by window duration, not by the old 5-word n-gram."""
self.assertEqual(
stitch_transcripts(
"well anyway one two three four five six",
"one two three four five six seven",
),
"well anyway one two three four five six seven",
)
def test_repeated_phrase_keeps_its_second_utterance(self):
"""Preferring the earliest match is what protects a real repeat."""
self.assertEqual(
stitch_transcripts("a b c fun time", "fun time fun time"),
"a b c fun time fun time",
)
def test_window_wholly_repeating_the_tail_is_dropped(self):
"""The accepted trade-off: an entirely redundant window adds nothing."""
self.assertEqual(stitch_transcripts("go go go", "go go go"), "go go go")
def test_revises_a_mistranscribed_tail(self):
"""A wrong last word would otherwise block every alignment.
Those words came from the newest audio, which the next window re-covers,
so replacing them is better than duplicating the phrase behind them.
"""
self.assertEqual(
stitch_transcripts("Yeah. this is Jessica.", "This is just gonna be fun."),
"Yeah. this is just gonna be fun.",
)
def test_revision_needs_more_than_one_shared_word(self):
"""A revision deletes published text, so it takes real evidence."""
self.assertEqual(
stitch_transcripts("the cat sat on a mat", "a dog barked"),
"the cat sat on a mat a dog barked",
)
if __name__ == "__main__":
unittest.main()
+6 -93
View File
@@ -12,7 +12,6 @@ from frigate.util.config import (
CURRENT_CONFIG_VERSION,
migrate_frigate_config,
migrate_models,
rename_hailo_detector,
)
@@ -129,10 +128,10 @@ class TestMigrateModels(unittest.TestCase):
migrated = migrate_models(
{
"detectors": {
"remote": {
"type": "zmq",
"endpoint": "tcp://host:5555",
"request_timeout_ms": 200,
"ds": {
"type": "deepstack",
"api_url": "http://host:5000/v1/vision/detection",
"api_key": "secret",
}
}
}
@@ -140,9 +139,9 @@ class TestMigrateModels(unittest.TestCase):
self.assertEqual(
migrated["models"][0]["devices"],
["zmq:tcp://host:5555"],
["deepstack:http://host:5000/v1/vision/detection"],
)
self.assertTrue(any("request_timeout_ms" in message for message in logs.output))
self.assertTrue(any("api_key" in message for message in logs.output))
def test_mixed_detector_types_are_logged(self):
with self.assertLogs("frigate.util.config", level=logging.ERROR) as logs:
@@ -165,61 +164,6 @@ class TestMigrateModels(unittest.TestCase):
self.assertEqual(migrated["mqtt"], {"host": "mqtt"})
class TestMigrateRenamedDetectors(unittest.TestCase):
def test_a_hailo_detector_is_renamed(self):
migrated = migrate_models(
{"detectors": {"hailo": {"type": "hailo8l", "device": "PCIe"}}}
)
self.assertEqual(migrated["models"][0]["devices"], ["hailo:PCIe"])
def test_a_hailo_detector_without_a_device(self):
migrated = migrate_models({"detectors": {"hailo": {"type": "hailo8l"}}})
self.assertEqual(migrated["models"][0]["devices"], ["hailo"])
def test_two_hailo_detectors_stay_separate(self):
# the shareable lookup has to resolve through the renamed key
migrated = migrate_models(
{
"detectors": {
"hailo1": {"type": "hailo8l", "device": "PCIe"},
"hailo2": {"type": "hailo8l", "device": "PCIe"},
}
}
)
self.assertEqual(migrated["models"][0]["devices"], ["hailo:PCIe", "hailo:PCIe"])
class TestRenameHailoDetector(unittest.TestCase):
def test_a_renamed_detector_is_updated(self):
migrated = rename_hailo_detector(
{"models": [{"devices": ["hailo8l:PCIe", "hailo8l"]}]}
)
self.assertEqual(migrated["models"][0]["devices"], ["hailo:PCIe", "hailo"])
def test_other_detectors_are_untouched(self):
migrated = rename_hailo_detector(
{"models": [{"devices": ["openvino:GPU", "cpu"]}]}
)
self.assertEqual(migrated["models"][0]["devices"], ["openvino:GPU", "cpu"])
def test_renaming_is_idempotent(self):
once = rename_hailo_detector({"models": [{"devices": ["hailo8l:PCIe"]}]})
twice = rename_hailo_detector(once)
self.assertEqual(twice["models"][0]["devices"], ["hailo:PCIe"])
def test_a_config_without_models_is_left_alone(self):
self.assertEqual(
rename_hailo_detector({"mqtt": {"enabled": False}}),
{"mqtt": {"enabled": False}},
)
class TestMigrateConfigFile(unittest.TestCase):
"""The full file migration, which is gated on shape as well as version."""
@@ -276,37 +220,6 @@ class TestMigrateConfigFile(unittest.TestCase):
os.path.exists(os.path.join(self.temp_dir.name, "backup_config.yaml"))
)
def test_a_legacy_hailo_detector_lands_on_the_hailo_key(self):
# the detectors key is folded into models after the version chain has
# run, so the rename has to happen there too
migrated = self._migrate(
"mqtt:\n"
" enabled: false\n"
"detectors:\n"
" hailo:\n"
" type: hailo8l\n"
" device: PCIe\n"
"cameras: {}\n"
"version: 0.18-0\n"
)
self.assertEqual(migrated["models"][0]["devices"], ["hailo:PCIe"])
self.assertNotIn("detectors", migrated)
def test_a_hailo8l_device_is_renamed_at_the_current_version(self):
migrated = self._migrate(
"mqtt:\n"
" enabled: false\n"
"models:\n"
" - scene: all\n"
" devices:\n"
" - hailo8l:PCIe\n"
"cameras: {}\n"
f"version: {CURRENT_CONFIG_VERSION}\n"
)
self.assertEqual(migrated["models"][0]["devices"], ["hailo:PCIe"])
if __name__ == "__main__":
unittest.main(verbosity=2)
+1 -48
View File
@@ -1,15 +1,10 @@
"""Tests for ONNX Runtime session option selection."""
import unittest
from unittest.mock import MagicMock, patch
import numpy as np
import onnxruntime as ort
from frigate.detectors.detection_runners import (
CudaGraphRunner,
get_ort_session_options,
)
from frigate.detectors.detection_runners import get_ort_session_options
from frigate.detectors.detector_config import ModelTypeEnum
from frigate.embeddings.types import EnrichmentModelTypeEnum
@@ -44,45 +39,3 @@ class TestGetOrtSessionOptions(unittest.TestCase):
]:
with self.subTest(model_type=model_type):
self.assertIsNone(get_ort_session_options(model_type))
class TestCudaGraphRunner(unittest.TestCase):
"""CUDA graph capture fails if the arena has to allocate during capture, so
the session is warmed up with capture disabled before the first real run."""
def setUp(self):
self.session = MagicMock()
self.session.get_outputs.return_value = [MagicMock(name="output")]
self.io_binding = self.session.io_binding.return_value
self.input = {"images": np.zeros((1, 3, 320, 320), np.float32)}
def _annotations(self) -> list[str | None]:
"""Graph annotation id passed with each run, None when unset."""
annotations = []
for call in self.session.run_with_iobinding.call_args_list:
try:
annotations.append(call.args[1].get_run_config_entry("gpu_graph_id"))
except RuntimeError:
annotations.append(None)
return annotations
def test_first_run_warms_up_with_capture_disabled(self):
with patch.object(ort.OrtValue, "ortvalue_from_numpy"):
CudaGraphRunner(self.session, 0).run(self.input)
self.assertEqual(
self._annotations(),
["-1"] * CudaGraphRunner.GRAPH_FREE_WARMUP_RUNS + [None],
)
def test_later_runs_allow_capture(self):
with patch.object(ort.OrtValue, "ortvalue_from_numpy"):
runner = CudaGraphRunner(self.session, 0)
runner.run(self.input)
self.session.run_with_iobinding.reset_mock()
runner.run(self.input)
self.assertEqual(self._annotations(), [None])
runner._input_ortvalue.update_inplace.assert_called_once()
+1 -1
View File
@@ -170,7 +170,7 @@ class TestAccelerators(HardwareProbeTestCase):
def test_hailo_is_found_by_its_device_node(self):
write(os.path.join(self.dev_root, "hailo0"))
self.assertEqual(self.probe()["hailo"].units[0].device, "hailo:PCIe")
self.assertEqual(self.probe()["hailo8l"].units[0].device, "hailo8l:PCIe")
def test_each_memryx_node_is_a_unit(self):
write(os.path.join(self.dev_root, "memx0"))
-54
View File
@@ -1,54 +0,0 @@
"""Tests for get_dst_transitions."""
import datetime
import unittest
from frigate.util.time import get_dst_transitions
class TestDstTransitions(unittest.TestCase):
def test_dst_transition_splits_periods_at_the_transition(self):
start = datetime.datetime(2026, 3, 7, 12, tzinfo=datetime.UTC).timestamp()
end = start + 2 * 86400
spring = datetime.datetime(2026, 3, 8, 7, tzinfo=datetime.UTC).timestamp()
self.assertEqual(
get_dst_transitions("America/New_York", start, end),
[(start, spring, -18000), (spring, end, -14400)],
)
def test_dst_transition_is_not_reported_a_day_late(self):
# local midnight on the day of the change used to report the
# transition a full day after it actually happened
start = datetime.datetime(2024, 3, 10, 5, tzinfo=datetime.UTC).timestamp()
end = start + 3 * 86400
spring = datetime.datetime(2024, 3, 10, 7, tzinfo=datetime.UTC).timestamp()
self.assertEqual(
get_dst_transitions("America/New_York", start, end),
[(start, spring, -18000), (spring, end, -14400)],
)
def test_dst_transition_after_the_last_daily_probe_is_found(self):
start = datetime.datetime(2024, 11, 2, 12, tzinfo=datetime.UTC).timestamp()
end = datetime.datetime(2024, 11, 3, 10, tzinfo=datetime.UTC).timestamp()
fall = datetime.datetime(2024, 11, 3, 6, tzinfo=datetime.UTC).timestamp()
self.assertEqual(
get_dst_transitions("America/New_York", start, end),
[(start, fall, -14400), (fall, end, -18000)],
)
def test_no_transition_returns_a_single_period(self):
start = datetime.datetime(2026, 6, 1, tzinfo=datetime.UTC).timestamp()
end = start + 5 * 86400
self.assertEqual(
get_dst_transitions("America/New_York", start, end),
[(start, end, -14400)],
)
def test_invalid_zone_retains_utc_fallback(self):
self.assertEqual(
get_dst_transitions("Invalid/Timezone", 100, 200), [(100, 200, 0)]
)
if __name__ == "__main__":
unittest.main(verbosity=2)
+3 -339
View File
@@ -360,70 +360,9 @@ class TestOllamaProvider(unittest.TestCase):
from frigate.genai.plugins.ollama import _normalize_multimodal_content
text, images = _normalize_multimodal_content(MULTIMODAL_MESSAGES[-1]["content"])
self.assertEqual(
text, "Here is the current live image from camera 'front'.\n[img]"
)
self.assertEqual(images, [b"\xff\xd8\xff\xd9"])
def test_normalize_keeps_text_and_image_order(self):
from frigate.genai.plugins.ollama import _normalize_multimodal_content
text, images = _normalize_multimodal_content(
[
{"type": "text", "text": "intro"},
{"type": "text", "text": "Frame 1"},
{"type": "image_url", "image_url": {"url": _IMAGE_DATA_URI}},
{"type": "text", "text": "Frame 2"},
{"type": "image_url", "image_url": {"url": _IMAGE_DATA_URI}},
]
)
self.assertEqual(text, "intro\nFrame 1\n[img]\nFrame 2\n[img]")
self.assertEqual(len(images), 2)
def test_send_uses_chat_with_captions_before_each_image(self):
client = self._client()
client.provider = MagicMock()
client.provider.chat.return_value = {
"message": {"content": '{"ok": true}'},
"done": True,
"done_reason": "stop",
}
client._supports_thinking_cache = False
result = client._send(
"prompt",
[b"a", b"b"],
{"type": "json_schema", "json_schema": {"schema": {"type": "object"}}},
image_captions=["Frame 1 of 2", "Frame 2 of 2"],
)
self.assertEqual(result, '{"ok": true}')
client.provider.generate.assert_not_called()
params = client.provider.chat.call_args.kwargs
self.assertEqual(
params["messages"],
[
{
"role": "user",
"content": "prompt\nFrame 1 of 2\n[img]\nFrame 2 of 2\n[img]",
"images": [b"a", b"b"],
}
],
)
self.assertEqual(params["format"], {"type": "object"})
self.assertNotIn("think", params)
def test_send_without_captions_puts_images_after_prompt(self):
client = self._client()
client.provider = MagicMock()
client.provider.chat.return_value = {"message": {"content": "ok"}, "done": True}
client._supports_thinking_cache = False
client._send("prompt", [b"a"])
message = client.provider.chat.call_args.kwargs["messages"][0]
self.assertEqual(message["content"], "prompt\n[img]")
self.assertEqual(message["images"], [b"a"])
self.assertIn("live image", text)
self.assertEqual(len(images), 1)
self.assertEqual(images[0], b"\xff\xd8\xff\xd9")
# ---------------------------------------------------------------------------
@@ -580,281 +519,6 @@ class TestLlamaCppProvider(unittest.TestCase):
client = self._validated_client(4096, {"context_size": 32768})
self.assertEqual(client.get_context_size(), 32768)
def test_list_models_dedupes_alias_matching_id(self):
client = self._client()
models_data = [
{"id": "qwen3-asr", "aliases": ["qwen3-asr"]},
{"id": "gemma", "aliases": ["gemma", "g4"]},
]
with patch.object(client, "_fetch_models_data", return_value=models_data):
self.assertEqual(client.list_models(), ["g4", "gemma", "qwen3-asr"])
# ---------------------------------------------------------------------------
# transcribe role
# ---------------------------------------------------------------------------
WAV_BYTES = b"RIFF$\x00\x00\x00WAVEfmt "
class TestOpenAITranscribe(unittest.TestCase):
def _client(self):
return _make_client(
"openai",
model="gpt-4o-transcribe",
api_key="k",
base_url="http://localhost:9999/v1",
runtime_options={"temperature": 0.7},
)
def test_supports_transcription(self):
self.assertTrue(self._client().supports_transcription)
def test_passes_file_tuple_and_language(self):
client = self._client()
create = MagicMock(return_value=" hello there ")
client.provider = SimpleNamespace(
audio=SimpleNamespace(transcriptions=SimpleNamespace(create=create))
)
self.assertEqual(client.transcribe(WAV_BYTES, language="en"), "hello there")
kwargs = create.call_args.kwargs
self.assertEqual(kwargs["model"], "gpt-4o-transcribe")
self.assertEqual(kwargs["file"], ("audio.wav", WAV_BYTES, "audio/wav"))
self.assertEqual(kwargs["language"], "en")
self.assertEqual(kwargs["response_format"], "text")
def test_does_not_forward_runtime_options(self):
"""runtime_options are chat parameters; /audio/transcriptions rejects them."""
client = self._client()
create = MagicMock(return_value="hi")
client.provider = SimpleNamespace(
audio=SimpleNamespace(transcriptions=SimpleNamespace(create=create))
)
client.transcribe(WAV_BYTES)
self.assertNotIn("temperature", create.call_args.kwargs)
def test_gpt_transcribe_uses_languages_array(self):
"""gpt-transcribe replaced `language` with a `languages` array."""
client = _make_client("openai", model="gpt-transcribe", api_key="k")
create = MagicMock(return_value="hi")
client.provider = SimpleNamespace(
audio=SimpleNamespace(transcriptions=SimpleNamespace(create=create))
)
client.transcribe(WAV_BYTES, language="en")
kwargs = create.call_args.kwargs
self.assertEqual(kwargs["extra_body"], {"languages": ["en"]})
# sending both fields is rejected by the API
self.assertNotIn("language", kwargs)
def test_older_models_use_singular_language(self):
for model in ("gpt-4o-transcribe", "whisper-1"):
with self.subTest(model):
client = _make_client("openai", model=model, api_key="k")
create = MagicMock(return_value="hi")
client.provider = SimpleNamespace(
audio=SimpleNamespace(transcriptions=SimpleNamespace(create=create))
)
client.transcribe(WAV_BYTES, language="en")
kwargs = create.call_args.kwargs
self.assertEqual(kwargs["language"], "en")
self.assertNotIn("extra_body", kwargs)
def test_object_response_form(self):
client = self._client()
create = MagicMock(return_value=SimpleNamespace(text="hi"))
client.provider = SimpleNamespace(
audio=SimpleNamespace(transcriptions=SimpleNamespace(create=create))
)
self.assertEqual(client.transcribe(WAV_BYTES), "hi")
def test_error_returns_none(self):
client = self._client()
create = MagicMock(side_effect=RuntimeError("boom"))
client.provider = SimpleNamespace(
audio=SimpleNamespace(transcriptions=SimpleNamespace(create=create))
)
self.assertIsNone(client.transcribe(WAV_BYTES))
class TestAzureOpenAITranscribe(unittest.TestCase):
def _client(self):
return _make_client(
"azure_openai",
model="my-deployment",
api_key="k",
base_url="https://example.openai.azure.com/?api-version=2024-06-01",
)
def test_routes_through_azure_client(self):
from openai import AzureOpenAI
client = self._client()
self.assertIsInstance(client.provider, AzureOpenAI)
self.assertTrue(client.supports_transcription)
def test_transcribe_inherited(self):
client = self._client()
create = MagicMock(return_value="azure text")
client.provider = SimpleNamespace(
audio=SimpleNamespace(transcriptions=SimpleNamespace(create=create))
)
self.assertEqual(client.transcribe(WAV_BYTES, language="fr"), "azure text")
self.assertEqual(create.call_args.kwargs["model"], "my-deployment")
class TestGeminiTranscribe(unittest.TestCase):
def _client(self):
return _make_client("gemini", model="gemini-2.0-flash", api_key="k")
def test_supports_transcription(self):
self.assertTrue(self._client().supports_transcription)
def test_sends_audio_part(self):
client = self._client()
generate = MagicMock(return_value=SimpleNamespace(text=" spoken words "))
client.provider = SimpleNamespace(
models=SimpleNamespace(generate_content=generate)
)
self.assertEqual(client.transcribe(WAV_BYTES, language="en"), "spoken words")
contents = generate.call_args.kwargs["contents"]
audio_parts = [
p for p in contents if getattr(p, "inline_data", None) is not None
]
self.assertEqual(len(audio_parts), 1)
self.assertEqual(audio_parts[0].inline_data.mime_type, "audio/wav")
self.assertEqual(audio_parts[0].inline_data.data, WAV_BYTES)
def test_oversized_payload_is_skipped(self):
from frigate.genai.plugins.gemini import GEMINI_MAX_INLINE_BYTES
client = self._client()
generate = MagicMock()
client.provider = SimpleNamespace(
models=SimpleNamespace(generate_content=generate)
)
self.assertIsNone(client.transcribe(b"\x00" * (GEMINI_MAX_INLINE_BYTES + 1)))
generate.assert_not_called()
class TestLlamaCppTranscribe(unittest.TestCase):
def _client(self, supports_audio: bool):
cfg = GenAIConfig(
provider="llamacpp",
model="m",
base_url="http://localhost:9999",
)
info = {
"context_size": 4096,
"supports_vision": False,
"supports_audio": supports_audio,
"supports_tools": False,
"supports_reasoning": False,
"media_marker": "<__media__>",
}
cls = PROVIDERS[GenAIProviderEnum.llamacpp]
with patch.object(cls, "_get_model_info", return_value=info):
return cls(cfg, timeout=5)
def test_supports_transcription_tracks_supports_audio(self):
self.assertTrue(self._client(True).supports_transcription)
self.assertFalse(self._client(False).supports_transcription)
@staticmethod
def _transcriptions_response(text: str = " transcript "):
response = MagicMock()
response.status_code = 200
response.json.return_value = {"text": text}
return response
@staticmethod
def _chat_response(content: str = " fallback transcript "):
response = MagicMock()
response.status_code = 200
response.json.return_value = {"choices": [{"message": {"content": content}}]}
return response
def test_posts_multipart_to_transcriptions(self):
client = self._client(True)
with patch.object(
client, "_post", return_value=self._transcriptions_response()
) as post:
self.assertEqual(client.transcribe(WAV_BYTES, language="en"), "transcript")
self.assertTrue(post.call_args.args[0].endswith("/v1/audio/transcriptions"))
self.assertEqual(
post.call_args.kwargs["files"]["file"],
("audio.wav", WAV_BYTES, "audio/wav"),
)
self.assertEqual(post.call_args.kwargs["data"]["language"], "en")
def test_omits_language_when_not_set(self):
"""An unset language is what lets the model detect one itself."""
client = self._client(True)
with patch.object(
client, "_post", return_value=self._transcriptions_response()
) as post:
client.transcribe(WAV_BYTES)
self.assertNotIn("language", post.call_args.kwargs["data"])
def test_falls_back_to_chat_completions_on_404(self):
"""Servers predating llama.cpp#21863 have no transcriptions route."""
client = self._client(True)
missing = MagicMock()
missing.status_code = 404
with patch.object(
client, "_post", side_effect=[missing, self._chat_response()]
) as post:
self.assertEqual(
client.transcribe(WAV_BYTES, language="en"), "fallback transcript"
)
urls = [call.args[0] for call in post.call_args_list]
self.assertTrue(urls[0].endswith("/v1/audio/transcriptions"))
self.assertTrue(urls[1].endswith("/v1/chat/completions"))
payload = post.call_args_list[1].kwargs["json"]
content = payload["messages"][0]["content"]
audio_parts = [p for p in content if p["type"] == "input_audio"]
self.assertEqual(len(audio_parts), 1)
self.assertEqual(audio_parts[0]["input_audio"]["format"], "wav")
self.assertEqual(
base64.b64decode(audio_parts[0]["input_audio"]["data"]), WAV_BYTES
)
def test_audio_unsupported_returns_none(self):
client = self._client(False)
with patch.object(client, "_post") as post:
self.assertIsNone(client.transcribe(WAV_BYTES))
post.assert_not_called()
class TestBaseClientTranscribe(unittest.TestCase):
"""Providers that don't implement the role must be inert, not broken."""
def test_ollama_reports_and_returns_nothing(self):
client = _make_client("ollama", model="llava", base_url="http://localhost:9999")
self.assertFalse(client.supports_transcription)
self.assertIsNone(client.transcribe(WAV_BYTES, language="en"))
if __name__ == "__main__":
unittest.main()
+2 -2
View File
@@ -200,7 +200,7 @@ class TestScanDetectors(HardwareStatsTestCase):
self.scan([DeviceSpec("openvino:CPU", "openvino", "CPU")]), set()
)
self.assertEqual(self.scan([DeviceSpec("cpu", "cpu", None)]), set())
self.assertEqual(self.scan([DeviceSpec("hailo", "hailo", None)]), set())
self.assertEqual(self.scan([DeviceSpec("hailo8l", "hailo8l", None)]), set())
def test_onnx_resolves_to_present_gpu(self):
spec = DeviceSpec("onnx", "onnx", None)
@@ -412,7 +412,7 @@ class TestHardwareTemperatures(unittest.TestCase):
return_value={"hailo8l-1": 52.0, "hailo8l-0": 51.0},
)
def test_hailo_sorted_by_name(self, temps):
self.assertEqual(get_hardware_temperatures("hailo"), [51.0, 52.0])
self.assertEqual(get_hardware_temperatures("hailo8l"), [51.0, 52.0])
def test_unsupported_type(self):
self.assertEqual(get_hardware_temperatures("rknn"), [])
-29
View File
@@ -73,35 +73,6 @@ class TestImprovedMotionDetector(unittest.TestCase):
"Motion boxes should be empty when scene change exceeds skip threshold",
)
def _bright_frame(self, offset: int) -> np.ndarray:
"""Produce a bright frame with a small dark object that moves."""
frame = np.full((self.frame_shape[0], self.frame_shape[1]), 200, dtype=np.uint8)
x = 10 + (offset * 5) % 60
frame[40:60, x : x + 15] = 20
return frame
def test_skip_motion_threshold_recovers_after_skip(self):
"""Skipped frames must still be blended into the background.
A bright scene differs from the zeroed background across the whole
frame, so the first frames are skipped. The background has to catch up
anyway, otherwise motion detection never returns.
"""
self.config.skip_motion_threshold = 0.5
self.config.improve_contrast = False
self.detector.config = self.config
self.detector.update_mask()
boxes = [len(self.detector.detect(self._bright_frame(i))) for i in range(40)]
self.assertEqual(boxes[0], 0, "First frame should exceed the skip threshold")
self.assertGreater(
self.detector.avg_frame.max(),
0,
"Background was never updated while frames were skipped",
)
self.assertTrue(any(boxes), "Motion detection never recovered after a skip")
def test_skip_motion_threshold_does_not_affect_calibration(self):
"""Even when skipping, the detector should go into calibrating state."""
self.config.skip_motion_threshold = 0.4
-79
View File
@@ -1,79 +0,0 @@
"""Tests for Frigate+ model metadata."""
import json
import os
import tempfile
import unittest
from unittest.mock import MagicMock, patch
from frigate.plus import PlusApi, load_plus_model_info
class TestHailoAlias(unittest.TestCase):
"""Frigate+ reports Hailo models by the detector's pre-rename key.
Both the model list and a cached info file have to carry the current key,
or every Hailo model reads as unsupported.
"""
def _get_models(self, models: list[dict]) -> list[dict]:
api = PlusApi.__new__(PlusApi)
response = MagicMock(ok=True, json=lambda: {"list": models})
with patch.object(PlusApi, "_get", return_value=response):
return api.get_models()["list"]
def test_the_model_list_gains_the_current_key(self):
models = self._get_models([{"supportedDetectors": ["hailo8l"]}])
self.assertEqual(models[0]["supportedDetectors"], ["hailo8l", "hailo"])
def test_the_old_key_is_kept_for_older_versions(self):
models = self._get_models([{"supportedDetectors": ["hailo8l"]}])
self.assertIn("hailo8l", models[0]["supportedDetectors"])
def test_other_detectors_are_untouched(self):
models = self._get_models([{"supportedDetectors": ["openvino", "onnx"]}])
self.assertEqual(models[0]["supportedDetectors"], ["openvino", "onnx"])
def test_a_model_already_naming_both_is_unchanged(self):
models = self._get_models([{"supportedDetectors": ["hailo8l", "hailo"]}])
self.assertEqual(models[0]["supportedDetectors"], ["hailo8l", "hailo"])
def test_an_empty_list(self):
self.assertEqual(self._get_models([]), [])
class TestLoadPlusModelInfo(unittest.TestCase):
def setUp(self):
self.cache = tempfile.TemporaryDirectory()
self.addCleanup(self.cache.cleanup)
patcher = patch("frigate.plus.MODEL_CACHE_DIR", self.cache.name)
patcher.start()
self.addCleanup(patcher.stop)
def _write(self, model_id: str, content: str) -> None:
with open(os.path.join(self.cache.name, f"{model_id}.json"), "w") as f:
f.write(content)
def test_a_cached_hailo_model_gains_the_current_key(self):
self._write("abc", json.dumps({"supportedDetectors": ["hailo8l"]}))
self.assertEqual(
load_plus_model_info("abc")["supportedDetectors"], ["hailo8l", "hailo"]
)
def test_a_missing_file(self):
self.assertIsNone(load_plus_model_info("nope"))
def test_an_unreadable_file(self):
self._write("bad", "{not json")
self.assertIsNone(load_plus_model_info("bad"))
if __name__ == "__main__":
unittest.main(verbosity=2)
-15
View File
@@ -75,21 +75,6 @@ class TestPreviewLoader(unittest.TestCase):
self.assertIsNone(get_most_recent_preview_frame(camera))
def test_get_most_recent_preview_frame_hyphenated_camera(self):
for name in ("preview_front-2000.0", "preview_front-door-3000.0"):
with open(
os.path.join(PREVIEW_CACHE_DIR, f"{name}.{PREVIEW_FRAME_TYPE}"), "w"
) as f:
f.write("test")
expected_path = os.path.join(
PREVIEW_CACHE_DIR, f"preview_front-2000.0.{PREVIEW_FRAME_TYPE}"
)
self.assertEqual(get_most_recent_preview_frame("front"), expected_path)
self.assertEqual(
get_most_recent_preview_frame("front", before=5000.0), expected_path
)
def test_get_most_recent_preview_frame_no_directory(self):
shutil.rmtree(PREVIEW_CACHE_DIR)
self.assertIsNone(get_most_recent_preview_frame("test_camera"))
-37
View File
@@ -15,8 +15,6 @@ Also covers the inverse direction: the ptz movement timestamps must not be writt
for a camera that has autotracking off, because nothing clears them back out.
"""
import asyncio
import threading
import unittest
from unittest.mock import AsyncMock, MagicMock
@@ -236,40 +234,5 @@ class TestManualRelativeMoveMetrics(unittest.IsolatedAsyncioTestCase):
)
class TestOnvifClose(unittest.TestCase):
"""close() must release everything on the loop, since whatever it leaves is
garbage collected during interpreter shutdown, where the resulting warnings
fail to log and fill the shutdown output with logging errors."""
def setUp(self) -> None:
self.controller = _make_controller(autotracking_enabled=False)
self.onvif = self.controller.cams[CAMERA]["onvif"]
self.onvif.close = AsyncMock()
self.controller.config_subscriber = MagicMock()
self.controller.loop = asyncio.new_event_loop()
self.controller.loop_thread = threading.Thread(
target=self.controller._run_event_loop, daemon=True
)
self.controller.loop_thread.start()
self.addCleanup(self.controller.loop.close)
def test_close_closes_camera_sessions(self) -> None:
self.controller.close()
self.onvif.close.assert_awaited_once()
def test_close_cancels_tasks_left_on_the_loop(self) -> None:
async def forever() -> None:
while True:
await asyncio.sleep(1)
poll = asyncio.run_coroutine_threadsafe(forever(), self.controller.loop)
self.controller.close()
self.assertTrue(poll.cancelled())
self.assertFalse(self.controller.loop_thread.is_alive())
if __name__ == "__main__":
unittest.main()
-55
View File
@@ -1,55 +0,0 @@
"""Tests for restarting frigate under s6."""
import signal
import unittest
from unittest.mock import MagicMock, patch
import psutil
from frigate.util.services import restart_frigate
class TestRestartFrigate(unittest.TestCase):
def _s6_process(self) -> MagicMock:
proc = MagicMock()
proc.name.return_value = "s6-svscan"
return proc
@patch("frigate.util.services.os.kill")
@patch("frigate.util.services.psutil.Process")
def test_terminates_s6_when_permitted(self, mock_process, mock_kill):
proc = self._s6_process()
mock_process.return_value = proc
restart_frigate()
proc.terminate.assert_called_once()
mock_kill.assert_not_called()
@patch("frigate.util.services.os.getpid", return_value=99)
@patch("frigate.util.services.os.kill")
@patch("frigate.util.services.psutil.Process")
def test_exits_self_when_s6_signal_is_denied(
self, mock_process, mock_kill, _mock_getpid
):
"""Running unprivileged, frigate cannot signal root's s6-svscan."""
proc = self._s6_process()
proc.terminate.side_effect = psutil.AccessDenied(pid=1, name="s6-svscan")
mock_process.return_value = proc
restart_frigate()
mock_kill.assert_called_once_with(99, signal.SIGINT)
@patch("frigate.util.services.os.getpid", return_value=99)
@patch("frigate.util.services.os.kill")
@patch("frigate.util.services.psutil.Process")
def test_exits_self_without_s6(self, mock_process, mock_kill, _mock_getpid):
proc = MagicMock()
proc.name.return_value = "init"
mock_process.return_value = proc
restart_frigate()
proc.terminate.assert_not_called()
mock_kill.assert_called_once_with(99, signal.SIGINT)
-316
View File
@@ -1,316 +0,0 @@
"""Tests for tracker-derived review frame annotations."""
import unittest
from frigate.data_processing.post.review_annotations import (
annotations_by_frame,
build_timeline,
describe_heading,
describe_position,
event_name,
path_legs,
path_moments,
)
def straight_path(
start: tuple[float, float],
end: tuple[float, float],
steps: int,
t0: float,
) -> list:
"""A path_data-shaped trajectory travelling in a straight line."""
return [
[
[
start[0] + (end[0] - start[0]) * i / (steps - 1),
start[1] + (end[1] - start[1]) * i / (steps - 1),
],
t0 + i,
]
for i in range(steps)
]
class TestDescribers(unittest.TestCase):
def test_describes_frame_corners(self):
self.assertEqual(describe_position(0.9, 0.9), "the bottom right of the frame")
self.assertEqual(describe_position(0.1, 0.1), "the top left of the frame")
self.assertEqual(describe_position(0.5, 0.5), "the middle of the frame")
def test_heading_treats_falling_y_as_up(self):
self.assertEqual(describe_heading(0.0, -0.5), "up")
self.assertEqual(describe_heading(0.0, 0.5), "down")
self.assertEqual(describe_heading(-0.5, 0.0), "left")
def test_heading_combines_axes(self):
self.assertEqual(describe_heading(-0.5, -0.5), "up and left")
class TestPathLegs(unittest.TestCase):
def test_straight_travel_is_one_leg(self):
points = [(0.9 - 0.05 * i, 0.6, float(i)) for i in range(10)]
self.assertEqual(len(path_legs(points)), 1)
def test_out_and_back_is_two_legs(self):
out = [(0.9 - 0.05 * i, 0.6, float(i)) for i in range(8)]
back = [(0.55 + 0.05 * i, 0.6, 8.0 + i) for i in range(8)]
self.assertEqual(len(path_legs(out + back)), 2)
def test_jitter_in_place_produces_no_legs(self):
# A subject standing still wobbles by a couple of percent; without the
# minimum leg distance this became a burst of contradictory turns.
points = [(0.5 + 0.01 * (i % 2), 0.5, float(i)) for i in range(20)]
self.assertEqual(path_legs(points), [])
def test_single_point_path_has_no_moments(self):
self.assertEqual(path_moments([[[0.5, 0.5], 1.0]]), [])
def test_malformed_path_is_ignored(self):
self.assertEqual(path_moments([[0.5, 1.0], [0.6, 2.0]]), [])
class TestEventNames(unittest.TestCase):
def test_unnamed_objects_use_an_indefinite_article(self):
self.assertEqual(event_name({"label": "person"}), "a person")
self.assertEqual(event_name({"label": "animal"}), "an animal")
def test_sub_labeled_objects_use_their_name(self):
event = {"label": "waste_bin", "sub_label": "Compost"}
self.assertEqual(event_name(event), 'waste bin "Compost"')
def test_repeated_objects_are_never_numbered(self):
# Whether these are the same subject is unknown, so the notes must not
# imply either answer.
events = [
{
"id": f"1789481994.68448{i}-abcdef",
"label": "person",
"sub_label": None,
"start_time": float(i * 10),
"end_time": float(i * 10 + 50),
"zones": [],
"path_data": straight_path((0.9, 0.6), (0.4, 0.3), 10, i * 10.0),
}
for i in range(3)
]
phrases = [p for _, p in build_timeline(events, span_end=100.0)]
self.assertTrue(all(p.startswith("a person ") for p in phrases))
self.assertFalse(any("#" in p for p in phrases))
class TestTimeline(unittest.TestCase):
def setUp(self):
self.event = {
"id": "1789481994.684479-lpyc2z",
"label": "person",
"sub_label": None,
"start_time": 0.0,
"end_time": 20.0,
"zones": ["front_yard"],
"path_data": straight_path((0.9, 0.6), (0.4, 0.3), 10, 1.0),
}
def test_timeline_is_ordered_and_bounded(self):
timeline = build_timeline([self.event], span_end=30.0)
times = [t for t, _ in timeline]
self.assertEqual(times, sorted(times))
self.assertTrue(all(t <= 30.0 for t in times))
def test_track_identifiers_never_appear(self):
timeline = build_timeline([self.event], span_end=30.0)
joined = " ".join(phrase for _, phrase in timeline)
self.assertNotIn("track", joined.lower())
self.assertNotIn(self.event["id"], joined)
def test_object_still_tracked_at_end_gets_no_closing_note(self):
# Its state at the end is visible in the last frame, so nothing is said.
timeline = build_timeline([self.event], span_end=15.0)
self.assertEqual(
[p for _, p in timeline],
["a person first detected at the right of the frame, moving up and left"],
)
def test_object_ending_inside_the_clip_is_no_longer_detected(self):
timeline = build_timeline([self.event], span_end=40.0)
phrases = [p for _, p in timeline]
self.assertTrue(any("no longer detected" in p for p in phrases))
def test_moments_past_the_last_frame_are_dropped(self):
# The subject keeps moving after the final sampled frame; those notes
# describe nothing the model can see.
late = dict(self.event)
out = straight_path((0.9, 0.6), (0.4, 0.6), 8, 100.0)
back = straight_path((0.4, 0.6), (0.9, 0.6), 8, 108.0)
late["path_data"] = out + back
late["start_time"] = 100.0
timeline = build_timeline([late], span_end=105.0)
self.assertTrue(any("moving left" in p for _, p in timeline))
self.assertFalse(any("turns around" in p for _, p in timeline))
def track(event_id, label, start, path, sub_label=None, end=None):
return {
"id": event_id,
"label": label,
"sub_label": sub_label,
"start_time": start,
"end_time": end if end is not None else start + 500.0,
"zones": [],
"path_data": path,
}
class TestArrival(unittest.TestCase):
def test_objects_arriving_together_keep_separate_notes(self):
person = track(
"1789481994.684479-lpyc2z",
"person",
0.0,
straight_path((0.9, 0.63), (0.58, 0.34), 10, 0.1),
)
bin_ = track(
"1789481995.063395-vlzd7q",
"waste_bin",
0.3,
straight_path((0.91, 0.64), (0.6, 0.33), 10, 0.4),
sub_label="Compost",
)
phrases = [p for _, p in build_timeline([person, bin_], 100.0)]
self.assertEqual(
phrases,
[
"a person first detected at the right of the frame, moving up and left",
'waste bin "Compost" first detected at the right of the frame, '
"moving up and left",
],
)
def test_object_detected_well_before_it_moves_gets_separate_notes(self):
# path_data always keeps the first two samples, so a bin sitting in
# the yard opens its leg long before it is picked up.
path = [[[0.91, 0.64], 0.4]] + straight_path(
(0.91, 0.64), (0.6, 0.33), 10, 20.4
)
bin_ = track(
"1789481995.063395-vlzd7q", "waste_bin", 0.3, path, sub_label="Compost"
)
timeline = build_timeline([bin_], 100.0)
self.assertEqual(
timeline[0],
(0.3, 'waste bin "Compost" first detected at the right of the frame'),
)
self.assertAlmostEqual(timeline[1][0], 21.4)
self.assertEqual(
timeline[1][1],
'waste bin "Compost" starts moving up and left from the right of the frame',
)
class TestStateChanges(unittest.TestCase):
def test_stationary_and_active_rows_become_notes(self):
bin_ = track(
"1789481995.063395-vlzd7q",
"waste_bin",
0.3,
straight_path((0.91, 0.64), (0.6, 0.33), 10, 0.4),
sub_label="Compost",
)
changes = [
{"timestamp": 20.0, "source_id": bin_["id"], "class_type": "stationary"},
{"timestamp": 30.0, "source_id": bin_["id"], "class_type": "active"},
{"timestamp": 35.0, "source_id": bin_["id"], "class_type": "entered_zone"},
{
"timestamp": 40.0,
"source_id": "someone-else",
"class_type": "stationary",
},
]
timeline = build_timeline([bin_], 100.0, state_changes=changes)
self.assertEqual(
timeline[1:],
[
(20.0, 'waste bin "Compost" has stopped moving'),
(30.0, 'waste bin "Compost" starts moving again'),
],
)
def test_state_changes_after_the_last_frame_are_dropped(self):
bin_ = track(
"1789481995.063395-vlzd7q",
"waste_bin",
0.3,
straight_path((0.91, 0.64), (0.6, 0.33), 10, 0.4),
sub_label="Compost",
)
changes = [
{"timestamp": 200.0, "source_id": bin_["id"], "class_type": "stationary"}
]
timeline = build_timeline([bin_], 100.0, state_changes=changes)
self.assertFalse(any("stopped" in p for _, p in timeline))
class TestNoAssumedState(unittest.TestCase):
def test_no_note_for_where_movement_ends(self):
# Ending a leftward walk still in the right third was read as a turn
# back to the right, and "stops moving" was read as standing still.
moments = path_moments(straight_path((0.95, 0.6), (0.7, 0.4), 6, 0.0))
self.assertEqual(
[p for _, p in moments],
["starts moving up and left from the right of the frame"],
)
def test_notes_never_claim_an_object_is_stationary(self):
# A track still open at the last frame says nothing about motion.
event = {
"id": "1789482056.695307-3uhf47",
"label": "person",
"sub_label": None,
"start_time": 0.0,
"end_time": 500.0,
"zones": ["front_yard"],
"path_data": straight_path((0.9, 0.6), (0.4, 0.3), 10, 1.0),
}
joined = " ".join(p for _, p in build_timeline([event], span_end=20.0))
self.assertNotIn("stationary", joined)
self.assertNotIn("stops", joined)
self.assertNotIn("leaves", joined)
self.assertNotIn("still", joined)
class TestLateEvents(unittest.TestCase):
def test_objects_first_detected_after_the_last_frame_are_skipped(self):
# Without this the arrival lands on the final frame, which was
# captured before the object appeared.
late = track(
"1789481999.000000-latear",
"person",
50.0,
straight_path((0.9, 0.6), (0.4, 0.3), 10, 50.1),
)
self.assertEqual(build_timeline([late], span_end=40.0), [])
self.assertEqual(
annotations_by_frame(build_timeline([late], 40.0), [0.0, 40.0]), {}
)
class TestFrameBucketing(unittest.TestCase):
def test_moment_attaches_to_the_following_frame(self):
frame_times = [0.0, 10.0, 20.0, 30.0]
buckets = annotations_by_frame([(12.0, "something happened")], frame_times)
self.assertEqual(buckets, {2: ["something happened"]})
def test_moment_on_a_frame_boundary_uses_that_frame(self):
buckets = annotations_by_frame([(10.0, "x")], [0.0, 10.0, 20.0])
self.assertEqual(buckets, {1: ["x"]})
def test_moment_after_the_last_frame_falls_on_the_last_frame(self):
buckets = annotations_by_frame([(99.0, "x")], [0.0, 10.0])
self.assertEqual(buckets, {1: ["x"]})
def test_no_frames_yields_no_buckets(self):
self.assertEqual(annotations_by_frame([(1.0, "x")], []), {})
if __name__ == "__main__":
unittest.main()
+3 -1
View File
@@ -324,7 +324,9 @@ class NorfairTracker(ObjectTracker):
):
tracker = self.get_tracker(obj["label"])
tracker.tracked_objects = [
o for o in tracker.tracked_objects if str(o.global_id) != track_id
o
for o in tracker.tracked_objects
if str(o.global_id) != track_id and o.hit_counter < 0
]
del self.track_id_map[track_id]
+7 -7
View File
@@ -406,12 +406,12 @@ class TrackedObjectProcessor(threading.Thread):
tracked_obj.obj_data["sub_label"] = (sub_label, score)
if event:
event.sub_label = sub_label
event.sub_label = sub_label # type: ignore[assignment]
data = event.data
if sub_label is None:
data["sub_label_score"] = None
data["sub_label_score"] = None # type: ignore[index]
elif score is not None:
data["sub_label_score"] = score
data["sub_label_score"] = score # type: ignore[index]
event.data = data
event.save()
@@ -440,7 +440,7 @@ class TrackedObjectProcessor(threading.Thread):
objects_list = []
sub_labels = set()
events = Event.select(Event.id, Event.label, Event.sub_label).where(
Event.id.in_(detection_ids)
Event.id.in_(detection_ids) # type: ignore[call-arg, misc]
)
for det_event in events:
if det_event.sub_label:
@@ -506,11 +506,11 @@ class TrackedObjectProcessor(threading.Thread):
if event:
data = event.data
data[field_name] = field_value
data[field_name] = field_value # type: ignore[index]
if field_value is None:
data[f"{field_name}_score"] = None
data[f"{field_name}_score"] = None # type: ignore[index]
elif score is not None:
data[f"{field_name}_score"] = score
data[f"{field_name}_score"] = score # type: ignore[index]
event.data = data
event.save()
+2 -263
View File
@@ -1,61 +1,16 @@
"""Utilities for creating and manipulating audio."""
import io
import logging
import os
import re
import string
import struct
import subprocess as sp
import wave
import numpy as np
from pathvalidate import sanitize_filename
from frigate.const import (
AUDIO_SAMPLE_RATE,
CACHE_DIR,
STREAM_TYPE_MAIN,
STREAM_TYPE_SUB,
)
from frigate.const import CACHE_DIR, STREAM_TYPE_MAIN, STREAM_TYPE_SUB
from frigate.models import Recordings
logger = logging.getLogger(__name__)
# Ceiling on the run of words the stitcher will treat as an overlap between two
# consecutive windows. This is an audio-duration bound, not a linguistic one: a
# window holds GENAI_WINDOW_CHUNKS * AUDIO_DURATION seconds of speech, so at a
# fast talker's pace it tops out around this many words, and a whole window can
# legitimately be redundant. The vendored whisper_streaming HypothesisBuffer
# caps at 5, but there the n-gram is only a tie-break on top of word-level
# timestamps; here it is the entire alignment, so 5 truncates real overlaps.
# Sentinel meaning "let the model work out the language". The vendored
# whisper_streaming code already uses this spelling, so it is the established
# convention for the audio_transcription.language field.
AUTO_LANGUAGE = "auto"
MAX_STITCH_NGRAM = 16
# How many trailing committed words the stitcher may discard to find an
# alignment. Those words came from the newest audio, which the next window
# re-covers, so when the provider got one of them wrong it blocks every
# alignment and the whole phrase duplicates. Set to 0 to make committed text
# strictly append-only.
MAX_STITCH_REVISE = 3
# A revision deletes text that was already published, so it has to clear a
# higher bar than a plain append: a single coincidentally shared word is not
# enough evidence to throw committed words away.
MIN_STITCH_REVISE_RUN = 2
# ASR models often wrap their output in control markup. Qwen3-ASR, for example,
# answers "language English<asr_text>Yeah, that works." A structural opening tag
# marks where the transcript starts, so anything before the last one is metadata.
# Closing tags (</x>) and pipe-delimited special tokens (<|endoftext|>) are
# excluded: those mark where the text ends, so text before them must be kept.
_OPENING_TAG = re.compile(r"<(?![/|])[^<>]*>")
_ANY_TAG = re.compile(r"<[^<>]*>")
def _get_recordings_for_range(
camera_name: str, start_ts: float, end_ts: float, stream_type: str
@@ -162,9 +117,7 @@ def get_audio_from_recording(
logger.debug(
f"Successfully extracted audio for {camera_name} from {start_ts} to {end_ts}"
)
# ffmpeg writes to a pipe, so it cannot seek back to patch the chunk
# sizes it reserved; repair them before any strict consumer sees them
return fix_wav_header(process.stdout)
return process.stdout
else:
logger.error(f"Failed to extract audio: {process.stderr.decode()}")
return None
@@ -176,217 +129,3 @@ def get_audio_from_recording(
os.unlink(file_path)
except OSError:
pass
def fix_wav_header(data: bytes) -> bytes:
"""Recompute the RIFF and data chunk sizes in a WAV header.
ffmpeg writing to a non-seekable pipe cannot go back and patch the sizes it
reserved, so it leaves 0xFFFFFFFF placeholders. PyAV-based demuxers ignore
them, but strict validators may reject the file or read zero frames.
Args:
data: The complete WAV payload
Returns:
The payload with both sizes corrected, or unchanged if it is not a
parseable RIFF/WAVE stream
"""
if len(data) < 12 or data[0:4] != b"RIFF" or data[8:12] != b"WAVE":
return data
out = bytearray(data)
# RIFF size covers everything after the 8-byte RIFF header
struct.pack_into("<I", out, 4, len(out) - 8)
# walk the chunk list to find "data"; every chunk is padded to even length
pos = 12
while pos + 8 <= len(out):
chunk_id = bytes(out[pos : pos + 4])
(chunk_size,) = struct.unpack_from("<I", out, pos + 4)
if chunk_id == b"data":
struct.pack_into("<I", out, pos + 4, len(out) - (pos + 8))
return bytes(out)
if chunk_size == 0xFFFFFFFF:
# an unpatched size before the data chunk leaves nothing to walk
break
pos += 8 + chunk_size + (chunk_size % 2)
return bytes(out)
def pcm16_to_wav(samples: np.ndarray, sample_rate: int = AUDIO_SAMPLE_RATE) -> bytes:
"""Wrap mono int16 PCM samples in a WAV container.
Args:
samples: The audio samples; converted to int16 if they are not already
sample_rate: Sample rate to declare in the header
Returns:
WAV bytes suitable for upload to a GenAI provider
"""
if samples.dtype != np.int16:
samples = samples.astype(np.int16)
buffer = io.BytesIO()
with wave.open(buffer, "wb") as wav:
wav.setnchannels(1)
wav.setsampwidth(2)
wav.setframerate(sample_rate)
wav.writeframes(samples.tobytes())
return buffer.getvalue()
def stitch_transcripts(committed: str, incoming: str) -> str:
"""Append *incoming* to *committed*, dropping the speech they share.
Consecutive overlapped transcription windows re-transcribe the same audio at
their seam, so the tail of one and the newest one name the same words. Find
the longest run that is a suffix of *committed* and occurs anywhere in
*incoming*, then keep only what follows that run.
Searching all of *incoming* rather than just its start is what makes this
work in practice. The provider re-transcribes the shared audio independently
and often gets its first word or two different ("just gonna" one window,
"It's gonna" the next), which leaves the real overlap sitting in the middle
of *incoming*. A prefix-anchored match sees no overlap at all there and
duplicates the entire phrase.
Text-level rather than timestamp-level because only some providers return
word timings, and this has to work across all of them.
Args:
committed: The transcript accumulated so far
incoming: The newest window's transcript
Returns:
The combined transcript
"""
incoming_words = incoming.split()
if not incoming_words:
return committed
committed_words = committed.split()
if not committed_words:
return " ".join(incoming_words)
committed_keys = [_overlap_key(word) for word in committed_words]
incoming_keys = [_overlap_key(word) for word in incoming_words]
length, consumed = _find_overlap(committed_keys, incoming_keys)
if length:
return " ".join(committed_words + incoming_words[consumed:])
# Nothing aligns. Retry against a shortened committed tail: a single word the
# provider got wrong at the end of the previous window otherwise blocks every
# alignment, and the entire re-transcribed phrase duplicates behind it.
best: tuple[int, int, int] | None = None
for drop in range(1, min(MAX_STITCH_REVISE, len(committed_keys) - 1) + 1):
length, consumed = _find_overlap(committed_keys[:-drop], incoming_keys)
if length < MIN_STITCH_REVISE_RUN:
continue
# longest run wins; ties go to the smallest revision
if best is None or length > best[0]:
best = (length, drop, consumed)
if best is None:
return " ".join(committed_words + incoming_words)
_, drop, consumed = best
return " ".join(committed_words[:-drop] + incoming_words[consumed:])
def _find_overlap(
committed_keys: list[str], incoming_keys: list[str]
) -> tuple[int, int]:
"""Locate the speech *incoming* shares with the end of *committed*.
Returns the length of the longest run that is a suffix of *committed_keys*
and occurs anywhere in *incoming_keys*, along with the index just past that
run in *incoming_keys*. Returns ``(0, 0)`` when nothing matches.
Prefers the longest run so a real overlap is not cut short, and within one
length the earliest position, so a phrase genuinely spoken twice keeps its
second utterance.
"""
max_run = min(MAX_STITCH_NGRAM, len(committed_keys), len(incoming_keys))
for length in range(max_run, 0, -1):
tail = committed_keys[-length:]
for start in range(len(incoming_keys) - length + 1):
if incoming_keys[start : start + length] == tail:
return length, start + length
return 0, 0
def clean_transcript(text: str | None) -> str:
"""Strip provider control markup and any preamble from a raw transcript.
A window with no speech often still comes back as the preamble alone
("language English<asr_text>"), which must reduce to an empty string so
callers treat it as silence rather than committing it as spoken words.
Args:
text: The provider's raw response
Returns:
The transcript with markup removed and whitespace collapsed
"""
if not text:
return ""
# everything up to and including the last opening tag is metadata
openings = list(_OPENING_TAG.finditer(text))
if openings:
text = text[openings[-1].end() :]
# drop closing tags and special tokens wherever they landed
text = _ANY_TAG.sub(" ", text)
return " ".join(text.split())
def _overlap_key(word: str) -> str:
"""Comparison key for overlap matching.
Providers re-transcribe the shared audio at a window seam independently, so
the same word routinely comes back capitalized differently or with different
edge punctuation ("work." vs "Work"). Those differences must not defeat the
match, but the original spelling is what gets kept in the output.
"""
key = word.strip(string.punctuation).casefold()
# a token that is nothing but punctuation would otherwise match any other
return key or word
def resolve_language(language: str | None) -> str | None:
"""Turn a configured language into an explicit code, or None for auto-detect.
Args:
language: The configured value, possibly AUTO_LANGUAGE
Returns:
An ISO language code, or None when the backend should detect it
"""
if not language or language == AUTO_LANGUAGE:
return None
return language
+4 -114
View File
@@ -31,6 +31,7 @@ DEFAULT_CONFIG_FILE = os.path.join(CONFIG_DIR, "config.yml")
DETECTOR_DEVICE_FIELDS = {
"cpu": "num_threads",
"rknn": "num_cores",
"deepstack": "api_url",
"degirum": "location",
"zmq": "endpoint",
}
@@ -39,6 +40,7 @@ DETECTOR_DEVICE_FIELDS = {
# detectors that use them are being reworked, so they are dropped rather than
# carried over.
DROPPED_DETECTOR_OPTIONS = {
"deepstack": ["api_timeout", "api_key"],
"degirum": ["zoo", "token"],
"zmq": ["request_timeout_ms", "linger_ms"],
}
@@ -168,20 +170,7 @@ def migrate_frigate_config(config_file: str):
# version and still use the pre-models detectors and model keys
needs_models = "detectors" in config or "model" in config
# likewise, it may already be on the models list and still name the hailo
# detector by its old key
needs_detector_rename = any(
isinstance(device, str) and device.partition(":")[0] == "hailo8l"
for model in (config.get("models") or [])
if isinstance(model, dict)
for device in (model.get("devices") or [])
)
if (
previous_version == CURRENT_CONFIG_VERSION
and not needs_models
and not needs_detector_rename
):
if previous_version == CURRENT_CONFIG_VERSION and not needs_models:
logger.info("frigate config does not need migration...")
return
@@ -255,12 +244,6 @@ def migrate_frigate_config(config_file: str):
with open(config_file, "w") as f:
yaml.dump(new_config, f)
if needs_detector_rename:
logger.info("Migrating renamed frigate detectors...")
new_config = rename_hailo_detector(new_config)
with open(config_file, "w") as f:
yaml.dump(new_config, f)
logger.info("Finished frigate config migration...")
@@ -794,95 +777,9 @@ def migrate_018_0(config: dict[str, dict[str, Any]]) -> dict[str, dict[str, Any]
return new_config
def rename_hailo_detector(
config: dict[str, dict[str, Any]],
) -> dict[str, dict[str, Any]]:
"""Rename the hailo8l detector, which drives every Hailo device.
Args:
config: The loaded config
Returns:
The config with every models entry naming the detector 'hailo'
"""
new_config = config.copy()
for model in new_config.get("models") or []:
if not isinstance(model, dict):
continue
devices = model.get("devices")
if not isinstance(devices, list):
continue
# assigned per index so ruamel keeps the comments on the list
for index, device in enumerate(devices):
if not isinstance(device, str):
continue
detector, separator, rest = device.partition(":")
if detector == "hailo8l":
devices[index] = f"hailo{separator}{rest}"
return new_config
def _camera_enables_transcription(camera: dict[str, Any]) -> bool:
"""Whether a camera or one of its profiles turns audio transcription on."""
sections = [camera.get("audio_transcription")]
profiles = camera.get("profiles")
if isinstance(profiles, dict):
for profile in profiles.values():
if isinstance(profile, dict):
sections.append(profile.get("audio_transcription"))
return any(
isinstance(section, dict) and section.get("enabled") for section in sections
)
def _migrate_transcription_language(config: dict[str, Any]) -> None:
"""Pin English for configs written before the language default became auto.
audio_transcription.language used to default to "en", so a config that
turned transcription on without naming a language was transcribing English.
The default is now "auto" (let the model detect), which is better for new
users but would silently change behavior for existing ones, so write the old
value explicitly for anyone actually using the feature.
"""
transcription = config.get("audio_transcription")
if isinstance(transcription, dict) and "language" in transcription:
# named a language already, so nothing was relying on the default
return
enabled = isinstance(transcription, dict) and bool(transcription.get("enabled"))
if not enabled:
enabled = any(
_camera_enables_transcription(camera)
for camera in config.get("cameras", {}).values()
if isinstance(camera, dict)
)
if not enabled:
return
if not isinstance(transcription, dict):
# a camera enabled it without a global section, which still picked up
# the global default
transcription = {}
config["audio_transcription"] = transcription
transcription["language"] = "en"
def migrate_019_0(config: dict[str, dict[str, Any]]) -> dict[str, dict[str, Any]]:
"""Handle migrating Frigate config to 0.19-0."""
new_config = rename_hailo_detector(config)
new_config = config.copy()
_migrate_birdseye_mode(new_config.get("birdseye"))
@@ -896,8 +793,6 @@ def migrate_019_0(config: dict[str, dict[str, Any]]) -> dict[str, dict[str, Any]
new_config["cameras"][name] = camera_config
_migrate_transcription_language(new_config)
new_config["version"] = "0.19-0"
return new_config
@@ -929,11 +824,6 @@ def migrate_models(config: dict[str, dict[str, Any]]) -> dict[str, dict[str, Any
detector = detector or {}
detector_type = detector.get("type", "cpu")
device = detector.get(DETECTOR_DEVICE_FIELDS.get(detector_type, "device"))
# hailo8l named one device, but the detector drives every Hailo device
if detector_type == "hailo8l":
detector_type = "hailo"
device_string = detector_type if device is None else f"{detector_type}:{device}"
# repeating a device now means running an extra inference process on it,
-3
View File
@@ -49,9 +49,6 @@ def start_or_restart_ffmpeg(
if ffmpeg_process is not None:
stop_ffmpeg(ffmpeg_process, logger)
# flush after the stop so the logs cover ffmpeg's output up to exit
logpipe.dump()
if frame_size is None:
process = sp.Popen(
ffmpeg_cmd,

Some files were not shown because too many files have changed in this diff Show More