Skip to content

Label Models

PhotoPrism's built-in image classification runs on ONNX Runtime and can be configured to use a different model than the one that is bundled. This page compares the available models, explains how to install and select one, and describes how developers and advanced users can plug in a custom classifier.

Development Preview

Configurable label models are available in our development preview builds and will be included in the next stable release. Current stable releases classify pictures with the TensorFlow model described in Using TensorFlow.

Built-in classifiers assign labels from a fixed vocabulary, run locally without network access, and need no GPU. If you want open-vocabulary labels or captions generated by a multimodal LLM instead, see Label Generation.

Available Models

All registered models were trained on ImageNet-1k, share the same 1,000-class vocabulary, and are mapped to the labels you see in the UI through the same label rules. Switching between them therefore changes how well pictures are recognized, not which label names can appear.

Name in vision.yml Model Installed by Default Download License Input Preprocessing
efficientformerv2_s2 EfficientFormerV2-S2 Yes 50.8 MB Apache-2.0 224 px, short edge 236, 95% crop, bicubic
efficientformerv2_s1 EfficientFormerV2-S1 No 24.8 MB Apache-2.0 224 px, short edge 236, 95% crop, bicubic
repvit_m1_0 RepViT-M1.0 No 27.4 MB Apache-2.0 224 px, short edge 236, 95% crop, bicubic
efficientnet_b0 EfficientNet-B0 RA No 21.2 MB Apache-2.0 224 px, short edge 256, 87.5% crop, bicubic

Every model is distributed as an FP32 graph with a pinned SHA-256 checksum. PhotoPrism verifies the checksum when it loads a registered model and refuses a file that does not match, rather than applying one model's preprocessing to another model's weights.

Comparison

The accuracy, parameter and compute figures below are published ImageNet-1k validation results for these checkpoints at 224 px. They describe the model on the ImageNet benchmark, not on personal photo collections:

Model Top-1 Top-5 Parameters GMACs Architecture
EfficientFormerV2-S2 82.2 95.9 12.7 M 1.25 Hybrid (convolution + attention)
RepViT-M1.0 80.4 94.9 7.3 M 1.10 Convolutional at inference
EfficientFormerV2-S1 79.7 94.7 6.2 M 0.65 Hybrid (convolution + attention)
EfficientNet-B0 RA 77.7 93.5 5.3 M 0.40 Convolutional

The following figures were measured by us. Each row names the hardware and method, because classification speed depends heavily on both:

Model Per Picture, p50 / p95 Measured On
EfficientFormerV2-S2 71 / 74 ms Intel Core i5-7500T, 4 cores, CPU provider; photoprism vision run -m labels, 48 sample images, warm
EfficientFormerV2-S1 68 / 72 ms same
RepViT-M1.0 76 / 81 ms same
EfficientNet-B0 RA 59 / 62 ms same

These times cover the complete labeling step for one picture, including loading and preparing its input. Note that EfficientFormerV2-S2 classifies up to two inputs per photo (see below), so its time usually includes two inferences while the other models perform one.

On a reviewed corpus of 402 Wikimedia photos, EfficientFormerV2-S2 assigned at least one visible label to 81.8% of the images, compared with 73.4% for EfficientFormerV2-S1, which left 107 rather than 73 images without a label. On ARM64, S2 measured 80 ms p50 and 121 ms p95 per inference with a peak memory usage of 350 MB, compared with 52 ms, 75 ms and 319 MB for S1. That higher cost was accepted in exchange for the better coverage when S2 was chosen as the default.

Choosing a Model

  • efficientformerv2_s2 is the default and the best choice for most libraries. It is the most accurate model and the only one that uses PhotoPrism's photo-specific input preparation: in addition to a center tile, it classifies the whole photo (cropping only extreme aspect ratios beyond 4:3), so subjects near the edges of wide photos can still be recognized.
  • efficientformerv2_s1 needs about half the disk space and less memory and CPU time, at the cost of more pictures without labels. Consider it for low-powered devices.
  • repvit_m1_0 is purely convolutional at inference, which tends to make its performance more predictable across CPU types.
  • efficientnet_b0 is the smallest and fastest model, and the most conservative choice in terms of operator support.

Alternative models classify a single center crop of each picture, as specified by their preprocessing, so content near the edges of panoramas and other wide images is not considered.

Labels are not recalibrated per model. The global confidence floor of 20% and the per-label thresholds in the label rules apply to every model, so the number of labels per picture can differ noticeably when you switch.

Selecting a Model

The following steps use repvit_m1_0 as an example and assume you are running our Docker image. Model files are installed below the models path, which is /opt/photoprism/assets/models in the image and can be changed with PHOTOPRISM_MODELS_PATH.

Step 1: Add a Model Folder

Since files written inside a container are lost when it is recreated, first create a folder for the model on your host and mount it into the container. Create the folder before starting the container so that it belongs to your user rather than to root:

mkdir -p models/repvit_m1_0

Then add a volume mount for it to the photoprism service in your compose.yaml and restart the service with docker compose up -d:

compose.yaml

services:
  photoprism:
    ...
    volumes:
      - "./models/repvit_m1_0:/opt/photoprism/assets/models/repvit_m1_0"
      ...

Mount only the subfolder of the model you are adding. Mounting an empty folder over /opt/photoprism/assets/models, or pointing PHOTOPRISM_MODELS_PATH to one, would hide the models that are installed by default, including the face detection and recognition models.

Step 2: Download the Model

Our images include a download script that fetches the model from dl.photoprism.app and verifies its checksum:

docker compose exec photoprism download-models.sh repvit_m1_0

The script is included in the PATH of our production images and the development environment, so it can be run without specifying a directory. Run download-models.sh --list to see all models it can install. It installs into the default models path and does not read PHOTOPRISM_MODELS_PATH, so if you have set a custom models path, pass the same directory as MODELS_PATH.

Step 3: Select the Model in vision.yml

Add a labels entry with the model name to the vision.yml file in your storage/config folder, or change the name of an existing labels entry:

vision.yml

Models:
  - Type: labels
    Name: repvit_m1_0
    Run: on-index

The default model runs while files are being indexed. Alternative models run in the background after indexing unless you set Run: on-index, as shown above, so newly indexed pictures receive their labels inline. The other Run values are described in the vision.yml reference.

A named model never falls back to different weights. If it cannot be loaded, no labels are generated until the problem is fixed and PhotoPrism is restarted, and a warning in the log tells you what went wrong.

Step 4: Restart & Verify

Restart PhotoPrism so that the model is loaded, then check the configuration:

docker compose restart photoprism
docker compose exec photoprism photoprism vision ls
docker compose exec photoprism photoprism vision status

vision ls should list your model with onnx as the engine, Enabled as the status, and Yes in the Installed column. vision status states which model generates labels and when, and shows the model file it loads:

Labels are generated by repvit_m1_0 on the cpu provider during indexing. ...

│ labels-model          │ auto (repvit_m1_0)                                         │
│ label-model-path      │ /opt/photoprism/assets/models/repvit_m1_0/repvit_m1_0.onnx │
│ label-model-runtime   │ onnx                                                       │

If the model file is missing, PhotoPrism logs a warning at startup stating that the label model is not installed, with the command to install it.

Step 5: Update Existing Labels

Changing the model does not modify labels that are already stored in your library. To classify existing pictures with the new model, run:

docker compose exec photoprism photoprism vision run -m labels --force

Use --count to try it on a limited number of pictures first, and --filter to select specific ones.

Reverting to the Default

To use the default model again, remove the labels entry from vision.yml or replace the model name with Default: true, then restart PhotoPrism. With Default: true, or no labels entry at all, PhotoPrism selects the first installed model in this order: efficientformerv2_s2, efficientformerv2_s1, repvit_m1_0, efficientnet_b0.

Disabling Classification

Set PHOTOPRISM_LABELS_MODEL to none to turn off built-in classification regardless of what vision.yml contains. The default value auto uses the model configured in vision.yml. Model names are not accepted by this option; they belong in vision.yml.

PHOTOPRISM_DISABLE_CLASSIFICATION is deprecated. It still applies when PHOTOPRISM_LABELS_MODEL is not set, but is ignored once PHOTOPRISM_LABELS_MODEL is set to auto or none.

Custom Models

You can also use your own image classifier, for example a model fine-tuned for a specific collection or one trained on a different vocabulary. Any labels entry whose name is not one of the registered models and that has no service endpoint is loaded as a custom ONNX model.

Requirements

  • One .onnx graph with a single output tensor of a fixed width that equals the number of entries in the label file.
  • A plain text label file with one class name per line, in output order. Lines must not be shifted by an extra background class: an output width that differs from the number of labels is rejected rather than producing labels that are off by one.
  • Raw logits or probabilities as output. Softmax is applied by PhotoPrism when ONNX.Output.Logits is true.
  • A fixed input size is recommended. For graphs with dynamic spatial axes, set the size with Resolution or ONNX.Input.

TensorFlow SavedModels are no longer supported. An existing labels entry for a TensorFlow model is replaced with the default model when vision.yml is loaded, and a warning is logged.

File Layout

By default, a model named my_classifier is loaded from these files below the models path:

my_classifier/my_classifier.onnx
my_classifier/labels.txt

Mount the folder as described in Step 1, e.g. ./models/my_classifier:/opt/photoprism/assets/models/my_classifier. Path changes the folder (relative to the models path), ONNX.File the file name of the graph, and LabelFile the name of the label file.

Configuration

Unlike the shape of the input and output tensors, which PhotoPrism reads from the graph, the color order, normalization and resize convention cannot be inferred from a model file. A model that is fed differently prepared images than it was trained on still loads and runs, but produces plausible-looking labels of noticeably lower quality. Declare these settings in vision.yml:

vision.yml

Models:
  - Type: labels
    Name: my_classifier
    Engine: onnx
    Run: on-index
    LabelFile: labels.txt
    ONNX:
      File: my_classifier.onnx
      Input:
        Width: 224
        Height: 224
        Layout: NCHW
        ColorOrder: RGB
        Normalization:
          Mean: [123.675, 116.28, 103.53]
          StdDev: [58.395, 57.12, 57.375]
        Resize:
          Mode: center-crop
          ShortEdge: 236
          CropRatio: 0.95
          Interpolation: bicubic
      Output:
        Logits: true
Field Values Description
ONNX.Input.Width / Height pixels Input size; read from the graph unless its spatial axes are dynamic
ONNX.Input.Layout NCHW, NHWC Tensor memory layout; read from the graph when possible
ONNX.Input.ColorOrder RGB, BGR Channel order expected by the model
ONNX.Input.Normalization Mean, StdDev Per-channel values on a 0-255 scale, in tensor channel order after ColorOrder is applied
ONNX.Input.Resize.Mode center-crop, stretch, pad How images are fitted to the input size
ONNX.Input.Resize.ShortEdge pixels Short edge the image is scaled to before a center crop
ONNX.Input.Resize.CropRatio 0-1 Crop ratio used to derive the short edge, e.g. 0.875 for 224 px out of 256 px
ONNX.Input.Resize.Interpolation nearest, linear, bicubic, lanczos Resampling filter
ONNX.Output.Logits true, false true if the output contains raw logits, false if it already contains probabilities
CanonicalOrder true, false Declares that the labels follow the canonical ImageNet-1k class order, which enables an additional check

Omitted preprocessing settings fall back to the ImageNet defaults shown in the example, and the assumed defaults are logged. An omitted Logits value is treated as true with a warning.

Instead of declaring these settings in vision.yml, they can also be embedded in the model file as photoprism.* metadata properties. Settings in vision.yml take precedence. Our export script scripts/ai/export-label-models.py shows how to export a timm checkpoint to ONNX with this metadata and verify the result against PyTorch: it exports a fixed [1, 3, 224, 224] FP32 graph at opset 17 with a single logits output and without a softmax layer.

Label Names

Class names are matched against the label rules in lowercase. Names that have a rule are renamed, assigned categories, or ignored as the rule specifies; names without a rule are used as they are once the model's confidence reaches 10% and the global confidence floor. This means that a model with a different vocabulary, such as ImageNet-21k, works without changes, but its class names may appear in the UI as written in your label file.

Labels generated by built-in and custom ONNX models are stored with the image source.

Label Rules

internal/ai/classify/rules.yml maps raw class names to the labels shown in the UI and sets per-label confidence thresholds, categories, and priorities:

cat:
  label: cat
  threshold: 0.3
  priority: 5
  categories:
    - animal

tabby cat:
  see: cat

After editing the rules, run go generate ./internal/ai/classify to regenerate rules.go and rebuild PhotoPrism. Rule changes only affect pictures that are classified afterwards.

Troubleshooting

Model Is Not Installed

vision ls shows No in the Installed column, and a startup warning names the download command. Make sure the model folder is mounted at the path shown by photoprism vision status (label-model-path), install the model as described in Step 2, and restart PhotoPrism. If the download fails with a permission error, make sure the user the container runs as can write to the mounted folder on your host.

Model Cannot Be Initialized

If a model fails to load, a warning such as init repvit_m1_0 model; fix or install it, then restart PhotoPrism is logged once and no labels are generated. The error is cached until PhotoPrism is restarted, so a repaired file is only picked up after a restart. Common causes are:

  • a registered model file whose checksum does not match, e.g. after an incomplete download (run the download script again with --force);
  • an output width that does not match the number of lines in the label file;
  • a graph with more than one output, or with an output width that is not fixed.

Labels Are Poor or Unexpected

For custom models, verify the preprocessing settings first, since a wrong channel order or normalization does not produce an error. To see which labels and confidence values a model returns for a few pictures, run:

docker compose exec photoprism photoprism --log-level=debug vision run -m labels --count 5 --force