omm

omm benchmark

Measure a small reproducible quality pack and decode speed for one or more installed models.

01 / 07

Overview

Reach for benchmark when you want real, comparable numbers instead of tune's prediction — it actually loads each model through Ollama (or LM Studio when Ollama isn't available) and runs the same fixed quality pack and repeated speed samples every time. Pass all to expand to everything installed for the active engine. Results are written as JSON evidence, and this is also what omm contribute runs in its loop.

02 / 07

Options

Every flag this command accepts, and what it defaults to when you leave it out.

  • <name>...—Default: required, or 'all'

    One or more installed models to benchmark, by Ollama tag or LM Studio modelKey — or the single word all to expand to everything installed.

  • --packPATHDefault: the built-in pack

    Use a different versioned quality pack instead of the built-in one.

  • --outputPATHDefault: an auto-generated path

    Write the evidence JSON to this path instead of an auto-generated one.

  • --speed-runs1-10Default: 3

    How many repeated speed samples to take before reporting a median.

  • --confirm-performance-timeout—Default: off

    If a model's first generation attempt times out, retry once instead of deciding immediately — see the flag's own help text for the exact tradeoff.

03 / 07

Examples

From a plain search to something you'd put in a script.

Benchmark one installed model.

$ omm benchmark qwen2.5-0.5b-instruct-q4_k_m.gguf

Benchmark every model installed for the active engine.

$ omm benchmark all

Take more speed samples for a steadier median.

$ omm benchmark qwen2.5-0.5b-instruct-q4_k_m.gguf --speed-runs 5

04 / 07

A real run

Real omm benchmark qwen2.5-0.5b-instruct-q4_k_m.gguf run, 2026-08-25, this dev machine — the real quality pack and speed samples actually ran through the real running Ollama. 1/8 (12.5%) is this 0.5B model's genuine score on the smoke pack, not a rosier invented one; 54.7 tok/s is this Apple M2's real measured decode speed for it. (This run's evidence was also really uploaded, since this dev machine already had upload policy set to always from earlier setup — the anonymized CPU/GPU score kind design/FACTS.md's setting section describes, nothing else.)

06 / 07

If something goes wrong

Every message below is one this command actually prints. Find yours, read why it happened, then do the last line.

  1. Neither Ollama nor LM Studio is installed or available. Install one of them, start it once, then retry `omm benchmark`.
    why
    Benchmarking needs a running engine to load the model into, and benchmark only knows how to drive Ollama or LM Studio for this.
    what to do
    Install and start Ollama or LM Studio at least once, then retry.
    source
    src/omm/cli.py:6825-6827
  2. `all` must be the only argument.
    why
    all expands to every installed model for the active engine — it can't be combined with other model names in the same run.
    what to do
    Pass all by itself, or list specific model names without it.
    source
    src/omm/cli.py:6820

Still stuck? Open an issue with the exact message you saw.

07 / 07

CLI reference

Exactly what omm benchmark --help prints, exported from the CLI source.

Usage

omm benchmark [OPTIONS] {models}...

Arguments

  • models...required

    One or more already-installed model identifiers for the active engine (Ollama tags, or LM Studio modelKeys when Ollama isn't available).

Options

  • --pack <path>Default: —

    Use a different versioned JSON pack.

  • --output <path>Default: —

    Write evidence to this JSON path.

  • --speed-runs <int range>Default: 3

    —

  • --confirm-performance-timeoutDefault: —

    If a model's first generation attempt times out, wait for it to fully finish, health-check the daemon, and retry exactly once before deciding. Two confirmed timeouts under a healthy daemon are reported as performance_unfit instead of transient_error. Off by default: a single timeout is never auto-retried unless you pass this flag.

Shared flags

Every omm command also accepts --json, --no-color, --quiet, -q, --yes, -y, so they are listed here once instead of on each command.

Exported from omm 0.3.101.

All commands