Model Card
A model's ID and instructions for use
A short, standardised document describing what a machine learning model was built for, how it was trained, how it performs on which tests, and where it should not be used.
Someone who downloads and uses a model doesn't know what the team that trained it knows: which data it saw, which groups it performs poorly on, which jobs it was never designed for. A model card is the document that carries that knowledge along with the model.
The concept comes from Margaret Mitchell et al., "Model Cards for Model Reporting" (arXiv 1810.03993, FAT* 2019). Proposed sections: - Model details: Developer, version, model type, licence. - Intended use: Primary uses and users; out-of-scope uses. - Factors: Groups, environments and instrumentation that can affect performance. - Metrics: Which measures, which thresholds, and why. - Evaluation data and training data. - Quantitative analyses: Results disaggregated by group. - Ethical considerations and caveats and recommendations.
Sibling concepts: - Datasheets for Datasets (Gebru et al., arXiv 1803.09010): The same idea for datasets; motivation, composition, collection process, preprocessing and labelling, uses, distribution and maintenance. - System card: A broader document published by large model providers that describes the model as part of a product, including safeguards, red-teaming findings and usage policies.
On Hugging Face, the model card is the repo's README.md. Its YAML
header (licence, language, base_model, datasets, evaluation results
under model-index) is machine-readable and feeds search and widgets;
the Markdown below is for humans.
On the regulatory side a model card is not a legal document on its own, but the content overlaps heavily. The EU AI Act requires high-risk system providers to keep technical documentation (Article 11, Annex IV), and requires general-purpose AI (GPAI) model providers to keep technical documentation and publish a summary of the content used for training following the AI Office template (Article 53(1)(d), applicable since 2 August 2025). A well-kept model card is the core of those documents.
Like the leaflet in a medicine box. Active ingredient, who can take it, dosage, side effects and "do not use during pregnancy", all on one page. The medicine might work without the leaflet; but when the wrong person takes the wrong dose, nobody knows who knew what and who is responsible.
A bank trains a Turkish BERT model that classifies the free-text notes on loan applications. The model card says:
- Intended use: Categorise applications to organise the analyst queue. Out of scope: approving or rejecting credit. - Training data: 180k anonymised application notes from 2023-2025; digital-channel applications are under-represented. - Evaluation: F1 0.87 on the test set, but 0.79 for applicants over 60. The gap is stated plainly. - Limitations: Performance drops on typo-heavy and code-mixed (Turkish-English) text.
Six months later another team wants to use the model for "automatic rejection". The out-of-scope line and the age-group gap on the card settle the discussion in the first meeting.
---
language:
- tr
license: apache-2.0
base_model: dbmdz/bert-base-turkish-cased
pipeline_tag: text-classification
tags:
- banking
- text-classification
model-index:
- name: loan-note-classifier
results:
- task:
type: text-classification
dataset:
name: Loan Application Notes 2026-Q2
type: bank/loan-notes # private dataset
split: test
metrics:
- type: f1
value: 0.87
---
# loan-note-classifier
## Intended use
Sorts application notes into 6 categories to organise the analyst queue.
### Out-of-scope use
- Approving, rejecting or scoring credit
- Drawing inferences about the applicant as a person
## Training data
180,000 anonymised notes from 2023-2025. Digital-channel
applications are under-represented (12%).
## Evaluation
| Group | F1 |
|----------------|------|
| Overall | 0.87 |
| Age 18-30 | 0.88 |
| Age 60+ | 0.79 |
## Limitations and risks
Performance drops on typo-heavy and Turkish-English mixed text.
Outputs must not be used without human review.- Whenever you share a model with another team, a customer or the public
- When maintaining an internal model inventory; every production model needs a card
- For high-risk or regulated use cases, as the backbone of technical documentation
- For fine-tuned models, to describe your changes on top of the base model's card
- Treating the card as full legal compliance; Annex IV technical documentation goes much further
- Writing a detailed card for throwaway experiment models nobody else will use
- Relying only on the provider's card for an API model; your use case isn't in it
Turning into marketing copy
A card that lists the best benchmark scores and skips the weaknesses defeats its own purpose. State disaggregated results and known failure modes plainly.
Written once, then stale
If the card isn't updated when the model is retrained, the data changes or a new bias finding appears, it becomes a source of misinformation. Generate it in the same CI step as the model version.
Leaving out-of-scope use empty
The out-of-scope section is the most useful part of the card. Leave it blank and the model quietly ends up making decisions it was never designed for.
Being vague about training data
'Web data' is not enough. List source types, date range, language mix and whether it contains personal data; for GPAI providers a public training-content summary is already a legal obligation.