REPRODUCIBLE MODEL TESTING

AI Image Quality Benchmarks: A Reproducible Test Method

Compare AI image models with fixed prompts, repeated outputs, blind reviews, weighted quality scores, and separate speed measurements.

Short answer

A fair AI image benchmark fixes the prompts, output count, settings, hardware, scoring rules, and review process before generation starts.

Four generated image samples arranged on a neutral evaluation grid with measurement tools

What does image quality mean in a benchmark?

Quality is a group of measured traits, not one universal visual score.

A useful benchmark separates prompt adherence, composition, visual defects, task usefulness, text rendering, and consistency.

T2I-CompBench shows why composition needs separate tests for attributes, objects, and relationships. HPS v2 shows that common automated metrics do not always match human preference.

Our view is simple: one attractive sample cannot prove that one model is better. The test must represent the work that a creator plans to do.

How large should the prompt set be?

Start with 24 prompts across six task groups and generate four outputs for every prompt.

This protocol produces 24 × 4 = 96 images for each model. Four models produce 384 review images.

Keep a separate unsupported-task result when a model cannot perform one task. Do not silently replace the task with an easier prompt.

  • Use four subject and portrait prompts.
  • Use four illustration and style prompts.
  • Use four spatial and multi-object composition prompts.
  • Use four readable-text and layout prompts.
  • Use four editing or image-guided prompts when every model supports them.
  • Use four normal product-workflow prompts from real creator needs.

How does the weighted quality score work?

Score each output from 0 to 5, then apply fixed weights that total 100.

For each category, multiply the 0-to-5 rating by its weight, then divide by 5. Add all six weighted results.

Example ratings of 4, 3, 4, 4, 2, and 3 produce 24 + 12 + 12 + 12 + 4 + 6 = 70 points.

Change the weights before testing when your workflow values another outcome. Never change them after seeing the images.

  • Prompt adherence: 30 points.
  • Composition and relationships: 20 points.
  • Visible defects: 15 points.
  • Task usefulness: 15 points.
  • Required text rendering: 10 points.
  • Consistency across four outputs: 10 points.

Which inputs must stay fixed?

Record every input that can change the result or the cost of producing it.

  • Record the exact model name, version, file, runtime, and license review date.
  • Record prompts, negative prompts, seeds, dimensions, steps, guidance, scheduler, and enhancement settings.
  • Record the device, system version, memory, storage, battery state, and starting temperature.
  • Record service plan, region, queue mode, and generation options for hosted models.
  • Keep every output, failure, retry, and moderation result in the test record.

How should speed and stability be compared?

Report speed, failures, and device effects separately from visual quality.

Use median generation time because one slow run can distort a simple average. Also report the 25th and 75th percentiles.

Calculate success rate as successful outputs ÷ attempted outputs × 100. For 93 successes from 96 attempts, the rate is 96.9%.

For local tests, report temperature change, battery change, peak memory when available, and throttling observations. For hosted tests, report queue time separately from processing time.

How do you reduce review bias?

Hide model names, randomize image order, and define rating examples before reviewers start.

  1. 01

    Create the protocol

    Freeze prompts, settings, weights, exclusions, and failure rules.

  2. 02

    Generate every output

    Keep failed and blocked attempts in the denominator.

  3. 03

    Blind the review

    Replace product names with random identifiers and shuffle the images.

  4. 04

    Check agreement

    Compare reviewer ratings and discuss categories with large differences.

  5. 05

    Publish the evidence

    Share prompts, settings, outputs, calculations, limitations, and review dates.

What can a benchmark conclusion prove?

It can support a decision for the tested conditions. It cannot prove universal model quality.

NIST measurement guidance emphasizes documented tests, metrics, processes, and sources of variation. That principle applies even when a benchmark uses human review.

A result can change with model versions, prompts, seeds, settings, devices, or reviewers. Publish these limits beside every score.

Offlair will use this method for future device and model tests. We will not publish model winners until complete outputs and provenance are available.

COMMON QUESTIONS

Frequently asked questions

How many images should an AI model benchmark include?

This starting method uses 96 images per model. Larger or risk-sensitive decisions need more prompts, repetitions, and reviewers.

Can one automated metric select the best image model?

No single metric covers prompt adherence, composition, preference, defects, usefulness, speed, and device fit.

Should local and cloud models use the same settings?

Use the same task and output goal. Record platform-specific settings because different systems do not expose identical controls.

Why keep failed generations in the results?

Failures affect real workflow reliability. Removing them can make an unstable model look better.

PRIMARY AND PRODUCT SOURCES

Sources and review notes

Product features, prices, and licenses can change. Review the linked source before a purchase or commercial use.