The systematic process of measuring how well an AI system performs against defined criteria.
Unlike traditional software testing, AI evaluation often deals with subjective quality (is this answer good?) rather than simple pass/fail checks, requiring different evaluation approaches like human review or LLM-as-judge.