Part XIII: Trust, Safety, Evaluation, and Operations
Chapter 65  [C]

Evaluation Protocols and Leakage-Safe Benchmarking

"A score I could not trust was worse than no score at all, so I learned to grade myself without cheating."

An Honestly Graded Embodied AI Agent

Placeholder chapter page. Sections are produced by the book-skills pipeline.

Sections

  1. 65.1 Task metrics vs operational metrics
  2. 65.2 User/site/device/session splits
  3. 65.3 Time-aware cross-validation
  4. 65.4 Rare-event and imbalanced evaluation
  5. 65.5 Robustness to missing and faulty sensors
  6. 65.6 Calibration and uncertainty metrics
  7. 65.7 Reproducibility and field-validation protocols

Lab 65

write an evaluation harness that reports standard metrics plus device/site/generalization breakdowns.