Box Commons

Comment on NIST AI 800-2: Evaluation Practices for Language Models

Date March 18, 2026
Submitted to National Institute of Standards and Technology
Docket NIST AI 800-2
Type Formal Comment (US Federal)

Box Commons · 30 N Gould St Ste N, Sheridan WY 82801

Key Takeaways
  • Evaluation results are already informing third-party credentialing and insurance underwriting — the stakes exceed academic benchmarking.
  • Behavioral safety is a distinct evaluation domain that NIST should formally recognize in AI 800-2.
  • Benchmark practices should address the needs of third-party credentialing use cases, not just model developers.
+ Jump to Section

I. Summary

NIST AI 800-2 provides a strong and timely framework for automated benchmark evaluation. The nine practices are well-structured, the agent evaluation coverage is substantive, and the document's candid acknowledgment that "automated benchmark evaluations cannot meet all AI evaluation objectives" reflects appropriate epistemic discipline.

We offer five observations where the document's framework could be strengthened, particularly as evaluation practices increasingly inform downstream decisions beyond internal development — including third-party credentialing, regulatory compliance, and insurance underwriting.

II. Behavioral Safety as a Distinct Evaluation Domain

The document addresses capabilities evaluation, security benchmarks (AgentDojo, CVE-Bench), and robustness testing, but does not recognize behavioral safety as a formalized evaluation domain.

An AI agent can pass every capabilities benchmark and security evaluation while still exhibiting unsafe behavioral patterns — escalating crisis interactions instead of deferring to human oversight, operating outside its delegated authority scope, or failing to disclose its non-human status in consumer-facing contexts. The judicial record confirms this distinction: in Garcia v. Character Technologies, the court found design defects in an agent's behavioral parameters, not its technical capabilities.

Recommendation: Practice 1.1 (Define Evaluation Objectives) should acknowledge behavioral safety as a distinct evaluation objective, separate from capabilities and security.

III. Evaluation for Third-Party Credentialing and Certification

The practices are framed exclusively around internal evaluation and voluntary assessment. The document does not address how these practices apply when evaluation results are used for downstream credentialing, certification, or conformance determinations.

Third-party evaluation for certification purposes introduces requirements that internal evaluation does not: reproducibility under adversarial conditions, defined pass/fail thresholds, chain of custody for evaluation artifacts, and evaluator independence. The document does not distinguish between self-evaluation and independent third-party evaluation — a distinction essential for credentialing legitimacy.

Recommendation: Include a discussion of how the nine practices apply differently when evaluation results are consumed for credentialing or certification purposes.

IV. Continuous and Periodic Re-Evaluation

The framework treats evaluation as a point-in-time activity. Language models and AI agents are not static artifacts — they receive updates, fine-tuning, and post-deployment alignment modifications. An evaluation conducted at deployment may not reflect the system's behavior six months later.

Recommendation: Add guidance on triggers for re-evaluation (model updates, fine-tuning, scaffolding changes), evaluation cadence for high-stakes applications, and how to design evaluation protocols efficient for repeated execution.

V. Downstream Consumers of Evaluation Results

Evaluation results are entering decision chains well beyond AI development teams: insurance underwriters are using evaluation data to price coverage, procurement officers are incorporating results into vendor selection, and state regulators in jurisdictions with AI oversight mandates may accept evaluation results as evidence of compliance.

Recommendation: Practice 3.3 (Report Qualified Claims) should address how to communicate evaluation results to non-technical downstream consumers — insurers, procurement officers, and regulators — who rely on them for consequential decisions.


Contact:
Brice Love, Acting Executive Director
Box Commons
[email protected]

Content Integrity Notice: This comment was authored by the Box Commons Policy Working Group. Generative AI was used for research synthesis and drafting support. All policy positions, recommendations, and normative claims were formulated and reviewed by human authors.