The Adversarial Platform Canvas
A reusable decision record for moving from ecosystem harm to proportionate intervention.
How to use it
Complete the canvas before selecting a detector. Cite evidence, mark assumptions, name an owner, and preserve alternatives. Revisit it after incidents, policy changes, major model changes, or attacker adaptation.
| # | Dimension | Required output |
|---|---|---|
| 1 | Ecosystem | Protected exchange, community, resource, or experience and its boundaries |
| 2 | Actors | Direct users, intermediaries, affected non-users, operators, and adversaries |
| 3 | Incentives | Benefits, costs, constraints, substitutes, and externalities for each actor |
| 4 | Harm | Observable undesirable outcomes, affected parties, magnitude, and reversibility |
| 5 | Policy | Allowed, limited, discouraged, restricted, and prohibited behavior |
| 6 | Observability | Necessary, lawful, reliable, and retained evidence with lineage |
| 7 | Detection | Inference methods and the claims each can actually support |
| 8 | Confidence | Calibration, base rate, manipulability, ambiguity, and disagreement |
| 9 | Urgency | Harm accumulation and the detection/decision latency budget |
| 10 | Intervention | Target, scope, severity, duration, reversibility, and expected effect |
| 11 | Error costs | False-positive, false-negative, friction, and collateral costs by population |
| 12 | Recourse | Notice, explanation, review, appeal, correction, and audit trail |
| 13 | Economics | Attacker and defender cost changes, scale limits, and displacement |
| 14 | Adaptation | Expected probing, evasion, imitation, migration, and next measurement |
Decision record
Decision:
Ecosystem and harm:
Policy authority:
Evidence and lineage:
Inference and confidence:
Latency budget:
Chosen intervention:
Why proportionate:
Error and subgroup risks:
Recourse and rollback:
Success and guardrail metrics:
Expected adaptation:
Owner and review date:
Quality checks
- Would the action still be justified if the actor were human rather than automated?
- Does the evidence support the intervention target, or only a related request/account/device?
- Is a less invasive or more reversible action sufficient?
- Can a legitimate user understand the rule and recover from error?
- What observation would prove the intervention ineffective or harmful?