What it is
A system card is a report that accompanies an AI model or product. It is the closest thing the industry has to a nutrition label. A good one tells you what the model is for, how it performed in testing, what risks the lab looked for, and where the model still falls short.
The name comes from "model cards," a documentation idea from machine learning research. System cards go further by covering the whole deployed system, including safeguards, rather than only the model's weights.
What is usually inside
Contents vary by lab, but a system card commonly covers:
- Capabilities. Results on an AI benchmark and other evaluations.
- Safety testing. What the lab tried in order to make the model misbehave, often including red teaming by internal staff or outside experts.
- Risk areas. Topics such as misuse for cyberattacks, dangerous biology or chemistry, deception and behavior as an AI agent.
- Mitigations. The AI guardrails and training steps applied, and what they did not fix.
- Known limits. Cases where the model is unreliable or where testing was incomplete.
Why it matters to you
If you build on a model, the system card is where a lab tells you what it did and did not check. You can use it to decide where to add your own tests, what to restrict in production, and which claims to repeat to your own customers.
Reading it well takes some skepticism. A system card is written by the lab that sells the model, so it shows what the lab chose to measure and report. Independent evaluations by outside auditors are a useful cross-check, because they test claims the lab cannot grade for itself. Notice also what is missing: a model released with no safety documentation tells you nothing about its testing, not that none happened.
Example
Before putting a new model in a customer support product, you open its system card, find the section on harmful content and prompt-based attacks, and note the weak spots. Then you write your own tests for those exact cases instead of assuming the lab covered your use.
System cards describe the results of work in AI alignment and often cite reasoning behavior such as a chain of thought. They differ from a benchmark, which measures one skill, because a system card covers risk and limits.