Independent reporting on artificial intelligence.


The Frontier Wire

Research

Reading a model card like an auditor

The model card was meant to be a nutrition label. Many have become brochures. A checklist for separating disclosure from marketing.

The model card was proposed in 2018 as a short, standardised document describing what a model is for, how it was evaluated, and where it should not be used. The idea caught on. Nearly every model release now ships with one. The quality has not kept pace with the adoption, and a card can now be anything from a rigorous disclosure to a benchmark table with a logo on it. This is a guide to reading one critically.

What a good card contains

The original proposal listed sections that remain the right checklist: intended use and out-of-scope uses, the evaluation data and metrics, disaggregated results across relevant groups, ethical considerations, and caveats. A card that covers each of those, with specifics, is doing its job.

The questions to ask

Which benchmarks, and which versions? Benchmark names are not enough. Many have multiple versions and evaluation settings, and scores are not comparable across them. A card that names a benchmark without a version or a prompt format is reporting a number you cannot check.

Who ran the evaluation? Vendor-run evaluations are the norm, and not necessarily wrong, but the card should say so. Independent results, where they exist, belong beside them.

Is anything disaggregated? An aggregate score hides variation. A card that reports one number per benchmark and nothing by language, domain or demographic group has not done the part of the work that model cards were invented for.

What is missing? The most informative section of many cards is the one that is not there. Training data is the usual gap. Compute is another. The absence is itself a disclosure.

A card that reports only the benchmarks the model wins is not a model card. It is a press release with a table.

Red flags

Be wary of cards that describe intended use so broadly that nothing is out of scope. Be wary of cards whose evaluation section changed between the preview and the general release without a changelog. And be wary of comparisons against competitors whose numbers were taken from a different evaluation setup than the vendor’s own.

Why it matters

Model cards are the document a buyer will be asked to produce when a regulator, a customer or an incident review asks what was known about the model at the time. Reading them like an auditor now is cheaper than explaining later why nobody did. For the same discipline applied to evaluation tools, see our guide to comparing evaluation frameworks.

Sources

  1. Model Cards for Model Reporting (Mitchell et al., 2018)
  2. Datasheets for Datasets (Gebru et al., 2018)
More research