Inspect AI
Model and agent evaluation harness · UK AI Security Institute · Intermediate to advanced
Rigorous, reproducible model and agent capability or safety evaluations involving tools, multi-turn interaction, coding, or sandboxed environments.
evals, llm evaluation, ai quality, open-source frameworks