Label Models¶
PhotoPrism's built-in image classification runs on ONNX Runtime and can be configured to use a different model than the one that is bundled. This page compares the available models, explains how to install and select one, and describes how developers and advanced users can plug in a custom classifier.
Development Preview
Configurable label models are available in our development preview builds and will be included in the next stable release. Current stable releases classify pictures with the TensorFlow model described in Using TensorFlow.
Built-in classifiers assign labels from a fixed vocabulary, run locally without network access, and need no GPU. If you want open-vocabulary labels or captions generated by a multimodal LLM instead, see Label Generation.
Available Models¶
All registered models were trained on ImageNet-1k, share the same 1,000-class vocabulary, and are mapped to the labels you see in the UI through the same label rules. Switching between them therefore changes how well pictures are recognized, not which label names can appear.
Name in vision.yml |
Model | Installed by Default | Download | License | Input Preprocessing |
|---|---|---|---|---|---|
efficientformerv2_s2 |
EfficientFormerV2-S2 | Yes | 50.8 MB | Apache-2.0 | 224 px, short edge 236, 95% crop, bicubic |
efficientformerv2_s1 |
EfficientFormerV2-S1 | No | 24.8 MB | Apache-2.0 | 224 px, short edge 236, 95% crop, bicubic |
repvit_m1_0 |
RepViT-M1.0 | No | 27.4 MB | Apache-2.0 | 224 px, short edge 236, 95% crop, bicubic |
efficientnet_b0 |
EfficientNet-B0 RA | No | 21.2 MB | Apache-2.0 | 224 px, short edge 256, 87.5% crop, bicubic |
Every model is distributed as an FP32 graph with a pinned SHA-256 checksum. PhotoPrism verifies the checksum when it loads a registered model and refuses a file that does not match, rather than applying one model's preprocessing to another model's weights.
Comparison¶
The accuracy, parameter and compute figures below are published ImageNet-1k validation results for these checkpoints at 224 px. They describe the model on the ImageNet benchmark, not on personal photo collections:
| Model | Top-1 | Top-5 | Parameters | GMACs | Architecture |
|---|---|---|---|---|---|
| EfficientFormerV2-S2 | 82.2 | 95.9 | 12.7 M | 1.25 | Hybrid (convolution + attention) |
| RepViT-M1.0 | 80.4 | 94.9 | 7.3 M | 1.10 | Convolutional at inference |
| EfficientFormerV2-S1 | 79.7 | 94.7 | 6.2 M | 0.65 | Hybrid (convolution + attention) |
| EfficientNet-B0 RA | 77.7 | 93.5 | 5.3 M | 0.40 | Convolutional |
The following figures were measured by us. Each row names the hardware and method, because classification speed depends heavily on both:
| Model | Per Picture, p50 / p95 | Measured On |
|---|---|---|
| EfficientFormerV2-S2 | 71 / 74 ms | Intel Core i5-7500T, 4 cores, CPU provider; photoprism vision run -m labels, 48 sample images, warm |
| EfficientFormerV2-S1 | 68 / 72 ms | same |
| RepViT-M1.0 | 76 / 81 ms | same |
| EfficientNet-B0 RA | 59 / 62 ms | same |
These times cover the complete labeling step for one picture, including loading and preparing its input. Note that EfficientFormerV2-S2 classifies up to two inputs per photo (see below), so its time usually includes two inferences while the other models perform one.
On a reviewed corpus of 402 Wikimedia photos, EfficientFormerV2-S2 assigned at least one visible label to 81.8% of the images, compared with 73.4% for EfficientFormerV2-S1, which left 107 rather than 73 images without a label. On ARM64, S2 measured 80 ms p50 and 121 ms p95 per inference with a peak memory usage of 350 MB, compared with 52 ms, 75 ms and 319 MB for S1. That higher cost was accepted in exchange for the better coverage when S2 was chosen as the default.
Choosing a Model¶
efficientformerv2_s2is the default and the best choice for most libraries. It is the most accurate model and the only one that uses PhotoPrism's photo-specific input preparation: in addition to a center tile, it classifies the whole photo (cropping only extreme aspect ratios beyond 4:3), so subjects near the edges of wide photos can still be recognized.efficientformerv2_s1needs about half the disk space and less memory and CPU time, at the cost of more pictures without labels. Consider it for low-powered devices.repvit_m1_0is purely convolutional at inference, which tends to make its performance more predictable across CPU types.efficientnet_b0is the smallest and fastest model, and the most conservative choice in terms of operator support.
Alternative models classify a single center crop of each picture, as specified by their preprocessing, so content near the edges of panoramas and other wide images is not considered.
Labels are not recalibrated per model. The global confidence floor of 20% and the per-label thresholds in the label rules apply to every model, so the number of labels per picture can differ noticeably when you switch.
Selecting a Model¶
The following steps use repvit_m1_0 as an example and assume you are running our Docker image. Model files are installed below the models path, which is /opt/photoprism/assets/models in the image and can be changed with PHOTOPRISM_MODELS_PATH.
Step 1: Add a Model Folder¶
Since files written inside a container are lost when it is recreated, first create a folder for the model on your host and mount it into the container. Create the folder before starting the container so that it belongs to your user rather than to root:
mkdir -p models/repvit_m1_0
Then add a volume mount for it to the photoprism service in your compose.yaml and restart the service with docker compose up -d:
compose.yaml
services:
photoprism:
...
volumes:
- "./models/repvit_m1_0:/opt/photoprism/assets/models/repvit_m1_0"
...
Mount only the subfolder of the model you are adding. Mounting an empty folder over /opt/photoprism/assets/models, or pointing PHOTOPRISM_MODELS_PATH to one, would hide the models that are installed by default, including the face detection and recognition models.
Step 2: Download the Model¶
Our images include a download script that fetches the model from dl.photoprism.app and verifies its checksum:
docker compose exec photoprism download-models.sh repvit_m1_0
The script is included in the PATH of our production images and the development environment, so it can be run without specifying a directory. Run download-models.sh --list to see all models it can install. It installs into the default models path and does not read PHOTOPRISM_MODELS_PATH, so if you have set a custom models path, pass the same directory as MODELS_PATH.
Step 3: Select the Model in vision.yml¶
Add a labels entry with the model name to the vision.yml file in your storage/config folder, or change the name of an existing labels entry:
vision.yml
Models:
- Type: labels
Name: repvit_m1_0
Run: on-index
The default model runs while files are being indexed. Alternative models run in the background after indexing unless you set Run: on-index, as shown above, so newly indexed pictures receive their labels inline. The other Run values are described in the vision.yml reference.
A named model never falls back to different weights. If it cannot be loaded, no labels are generated until the problem is fixed and PhotoPrism is restarted, and a warning in the log tells you what went wrong.
Step 4: Restart & Verify¶
Restart PhotoPrism so that the model is loaded, then check the configuration:
docker compose restart photoprism
docker compose exec photoprism photoprism vision ls
docker compose exec photoprism photoprism vision status
vision ls should list your model with onnx as the engine, Enabled as the status, and Yes in the Installed column. vision status states which model generates labels and when, and shows the model file it loads:
Labels are generated by repvit_m1_0 on the cpu provider during indexing. ...
│ labels-model │ auto (repvit_m1_0) │
│ label-model-path │ /opt/photoprism/assets/models/repvit_m1_0/repvit_m1_0.onnx │
│ label-model-runtime │ onnx │
If the model file is missing, PhotoPrism logs a warning at startup stating that the label model is not installed, with the command to install it.
Step 5: Update Existing Labels¶
Changing the model does not modify labels that are already stored in your library. To classify existing pictures with the new model, run:
docker compose exec photoprism photoprism vision run -m labels --force
Use --count to try it on a limited number of pictures first, and --filter to select specific ones.
Reverting to the Default¶
To use the default model again, remove the labels entry from vision.yml or replace the model name with Default: true, then restart PhotoPrism. With Default: true, or no labels entry at all, PhotoPrism selects the first installed model in this order: efficientformerv2_s2, efficientformerv2_s1, repvit_m1_0, efficientnet_b0.
Disabling Classification¶
Set PHOTOPRISM_LABELS_MODEL to none to turn off built-in classification regardless of what vision.yml contains. The default value auto uses the model configured in vision.yml. Model names are not accepted by this option; they belong in vision.yml.
PHOTOPRISM_DISABLE_CLASSIFICATION is deprecated. It still applies when PHOTOPRISM_LABELS_MODEL is not set, but is ignored once PHOTOPRISM_LABELS_MODEL is set to auto or none.
Custom Models¶
You can also use your own image classifier, for example a model fine-tuned for a specific collection or one trained on a different vocabulary. Any labels entry whose name is not one of the registered models and that has no service endpoint is loaded as a custom ONNX model.
Requirements¶
- One
.onnxgraph with a single output tensor of a fixed width that equals the number of entries in the label file. - A plain text label file with one class name per line, in output order. Lines must not be shifted by an extra background class: an output width that differs from the number of labels is rejected rather than producing labels that are off by one.
- Raw logits or probabilities as output. Softmax is applied by PhotoPrism when
ONNX.Output.Logitsistrue. - A fixed input size is recommended. For graphs with dynamic spatial axes, set the size with
ResolutionorONNX.Input.
TensorFlow SavedModels are no longer supported. An existing labels entry for a TensorFlow model is replaced with the default model when vision.yml is loaded, and a warning is logged.
File Layout¶
By default, a model named my_classifier is loaded from these files below the models path:
my_classifier/my_classifier.onnx
my_classifier/labels.txt
Mount the folder as described in Step 1, e.g. ./models/my_classifier:/opt/photoprism/assets/models/my_classifier. Path changes the folder (relative to the models path), ONNX.File the file name of the graph, and LabelFile the name of the label file.
Configuration¶
Unlike the shape of the input and output tensors, which PhotoPrism reads from the graph, the color order, normalization and resize convention cannot be inferred from a model file. A model that is fed differently prepared images than it was trained on still loads and runs, but produces plausible-looking labels of noticeably lower quality. Declare these settings in vision.yml:
vision.yml
Models:
- Type: labels
Name: my_classifier
Engine: onnx
Run: on-index
LabelFile: labels.txt
ONNX:
File: my_classifier.onnx
Input:
Width: 224
Height: 224
Layout: NCHW
ColorOrder: RGB
Normalization:
Mean: [123.675, 116.28, 103.53]
StdDev: [58.395, 57.12, 57.375]
Resize:
Mode: center-crop
ShortEdge: 236
CropRatio: 0.95
Interpolation: bicubic
Output:
Logits: true
| Field | Values | Description |
|---|---|---|
ONNX.Input.Width / Height |
pixels | Input size; read from the graph unless its spatial axes are dynamic |
ONNX.Input.Layout |
NCHW, NHWC |
Tensor memory layout; read from the graph when possible |
ONNX.Input.ColorOrder |
RGB, BGR |
Channel order expected by the model |
ONNX.Input.Normalization |
Mean, StdDev |
Per-channel values on a 0-255 scale, in tensor channel order after ColorOrder is applied |
ONNX.Input.Resize.Mode |
center-crop, stretch, pad |
How images are fitted to the input size |
ONNX.Input.Resize.ShortEdge |
pixels | Short edge the image is scaled to before a center crop |
ONNX.Input.Resize.CropRatio |
0-1 |
Crop ratio used to derive the short edge, e.g. 0.875 for 224 px out of 256 px |
ONNX.Input.Resize.Interpolation |
nearest, linear, bicubic, lanczos |
Resampling filter |
ONNX.Output.Logits |
true, false |
true if the output contains raw logits, false if it already contains probabilities |
CanonicalOrder |
true, false |
Declares that the labels follow the canonical ImageNet-1k class order, which enables an additional check |
Omitted preprocessing settings fall back to the ImageNet defaults shown in the example, and the assumed defaults are logged. An omitted Logits value is treated as true with a warning.
Instead of declaring these settings in vision.yml, they can also be embedded in the model file as photoprism.* metadata properties. Settings in vision.yml take precedence. Our export script scripts/ai/export-label-models.py shows how to export a timm checkpoint to ONNX with this metadata and verify the result against PyTorch: it exports a fixed [1, 3, 224, 224] FP32 graph at opset 17 with a single logits output and without a softmax layer.
Label Names¶
Class names are matched against the label rules in lowercase. Names that have a rule are renamed, assigned categories, or ignored as the rule specifies; names without a rule are used as they are once the model's confidence reaches 10% and the global confidence floor. This means that a model with a different vocabulary, such as ImageNet-21k, works without changes, but its class names may appear in the UI as written in your label file.
Labels generated by built-in and custom ONNX models are stored with the image source.
Label Rules¶
internal/ai/classify/rules.yml maps raw class names to the labels shown in the UI and sets per-label confidence thresholds, categories, and priorities:
cat:
label: cat
threshold: 0.3
priority: 5
categories:
- animal
tabby cat:
see: cat
After editing the rules, run go generate ./internal/ai/classify to regenerate rules.go and rebuild PhotoPrism. Rule changes only affect pictures that are classified afterwards.
Troubleshooting¶
Model Is Not Installed¶
vision ls shows No in the Installed column, and a startup warning names the download command. Make sure the model folder is mounted at the path shown by photoprism vision status (label-model-path), install the model as described in Step 2, and restart PhotoPrism. If the download fails with a permission error, make sure the user the container runs as can write to the mounted folder on your host.
Model Cannot Be Initialized¶
If a model fails to load, a warning such as init repvit_m1_0 model; fix or install it, then restart PhotoPrism is logged once and no labels are generated. The error is cached until PhotoPrism is restarted, so a repaired file is only picked up after a restart. Common causes are:
- a registered model file whose checksum does not match, e.g. after an incomplete download (run the download script again with
--force); - an output width that does not match the number of lines in the label file;
- a graph with more than one output, or with an output width that is not fixed.
Labels Are Poor or Unexpected¶
For custom models, verify the preprocessing settings first, since a wrong channel order or normalization does not produce an error. To see which labels and confidence values a model returns for a few pictures, run:
docker compose exec photoprism photoprism --log-level=debug vision run -m labels --count 5 --force