Comment on NIST AI 800-2: Evaluation Practices for Language Models
Box Commons · 30 N Gould St Ste N, Sheridan WY 82801
- Evaluation results are already informing third-party credentialing and insurance underwriting — the stakes exceed academic benchmarking.
- Behavioral safety is a distinct evaluation domain that NIST should formally recognize in AI 800-2.
- Benchmark practices should address the needs of third-party credentialing use cases, not just model developers.
+ Jump to Section
I. Summary
NIST AI 800-2 provides a strong and timely framework for automated benchmark evaluation. The nine practices are well-structured, the agent evaluation coverage is substantive, and the document's candid acknowledgment that "automated benchmark evaluations cannot meet all AI evaluation objectives" reflects appropriate epistemic discipline.
We offer five observations where the document's framework could be strengthened, particularly as evaluation practices increasingly inform downstream decisions beyond internal development — including third-party credentialing, regulatory compliance, and insurance underwriting.
II. Behavioral Safety as a Distinct Evaluation Domain
The document addresses capabilities evaluation, security benchmarks (AgentDojo, CVE-Bench), and robustness testing, but does not recognize behavioral safety as a formalized evaluation domain.
An AI agent can pass every capabilities benchmark and security evaluation while still exhibiting unsafe behavioral patterns — escalating crisis interactions instead of deferring to human oversight, operating outside its delegated authority scope, or failing to disclose its non-human status in consumer-facing contexts. The judicial record confirms this distinction: in Garcia v. Character Technologies, the court found design defects in an agent's behavioral parameters, not its technical capabilities.
Recommendation: Practice 1.1 (Define Evaluation Objectives) should acknowledge behavioral safety as a distinct evaluation objective, separate from capabilities and security.
III. Evaluation for Third-Party Credentialing and Certification
The practices are framed exclusively around internal evaluation and voluntary assessment. The document does not address how these practices apply when evaluation results are used for downstream credentialing, certification, or conformance determinations.
Third-party evaluation for certification purposes introduces requirements that internal evaluation does not: reproducibility under adversarial conditions, defined pass/fail thresholds, chain of custody for evaluation artifacts, and evaluator independence. The document does not distinguish between self-evaluation and independent third-party evaluation — a distinction essential for credentialing legitimacy.
Recommendation: Include a discussion of how the nine practices apply differently when evaluation results are consumed for credentialing or certification purposes.
IV. Continuous and Periodic Re-Evaluation
The framework treats evaluation as a point-in-time activity. Language models and AI agents are not static artifacts — they receive updates, fine-tuning, and post-deployment alignment modifications. An evaluation conducted at deployment may not reflect the system's behavior six months later.
Recommendation: Add guidance on triggers for re-evaluation (model updates, fine-tuning, scaffolding changes), evaluation cadence for high-stakes applications, and how to design evaluation protocols efficient for repeated execution.
V. Downstream Consumers of Evaluation Results
Evaluation results are entering decision chains well beyond AI development teams: insurance underwriters are using evaluation data to price coverage, procurement officers are incorporating results into vendor selection, and state regulators in jurisdictions with AI oversight mandates may accept evaluation results as evidence of compliance.
Recommendation: Practice 3.3 (Report Qualified Claims) should address how to communicate evaluation results to non-technical downstream consumers — insurers, procurement officers, and regulators — who rely on them for consequential decisions.
Contact:
Brice Love, Acting Executive Director
Box Commons
[email protected]
Content Integrity Notice: This comment was authored by the Box Commons Policy Working Group. Generative AI was used for research synthesis and drafting support. All policy positions, recommendations, and normative claims were formulated and reviewed by human authors.
Related Filings
Comment on NIST CAISI RFI 2025-0035: AI Agent Security
Response to the Center for AI Safety and Innovation's request for information on AI agent security standards. Argues that behavioral safety is a distinct, unaddressed security domain and that NIST should expand 'agent security' to include a behavioral safety layer suitable for insurance underwriting.
NISTComment on NCCoE AI Agent Identity and Authorization
Recommends that agent identity credentials be interlocked with behavioral safety verification — no credential without verified behavioral certification. Identity without behavioral verification provides incomplete assurance to downstream systems.