waza Skill
WORKFLOW SKILL - Evaluate AI agent skills using structured benchmarks with YAML specs, fixture isolation, and pluggable validators. USE FOR: run waza, waza help, run eval, run benchmark, evaluate skill, test agent, generate eval suite, init eval, compare results, score agent, agent evaluation, skill testing, cross-mode Published by microsoft in waza.
What is waza Skill?
WORKFLOW SKILL - Evaluate AI agent skills using structured benchmarks with YAML specs, fixture isolation, and pluggable validators. USE FOR: run waza, waza help, run eval, run benchmark, evaluate skill, test agent, generate eval suite, init eval, compare results, score agent, agent evaluation, skill testing, cross-mode Published by microsoft in waza. This profile combines repository metadata with install, compatibility, and usage signals so developers can quickly decide whether it fits their agent workflow before opening the source repository.
Automated repository signals based on public metadata such as recency, license, installation evidence, and adoption. These are not a security audit or endorsement.
Key capabilities
- Includes SKILL.md support
- Reusable instructions support
- Testing
- Testing use cases
Technical details
- Install or run with Copy skill directory
When to use waza Skill
- Use it for testing.
Built with
Editorial notes
Source
- Creator: microsoft
- Repository: microsoft/waza
- Skill file: skills/waza/SKILL.md
What it does
WORKFLOW SKILL - Evaluate AI agent skills using structured benchmarks with YAML specs, fixture isolation, and pluggable validators. USE FOR: run waza, waza help, run eval, run benchmark, evaluate skill, test agent, generate eval suite, init eval, compare results, score agent, agent evaluation, skill testing, cross-mode
Skill instructions
Waza "The way of technique — measure, refine, master." A Go CLI tool for evaluating AI agent skills through structured benchmarks. Define test cases in YAML, run them against agent engines, and validate results with pluggable scoring validators. Help When user says "waza help" or asks how to use waza: ╔══════════════════════════════════════════════════════════════════╗ ║ WAZA - CLI Tool for Evaluating Agent Skills ║ ╠══════════════════════════════════════════════════════════════════╣ ║ ║ ║ COMMANDS: ║ ║ waza run <eval.yaml Run an evaluation benchmark ║ ║ waza init [directory] Initialize a new eval suite ║ ║ waza generate <SKILL.md Generate eval from SKILL.md ║ ║ waza compare <r1 <r2 ... Compare result files ║ ║ waza dev [skill-path] Improve SKILL.md compliance ║ ║ ║ ║ RUN FLAGS: ║ ║ --context-dir, -c Fixtures directory (default: ./fixtures) ║ ║ --output, -o Save results JSON to file ║ ║ --verbose, -v Verbose output ║ ║ --task, -t Filter tasks by name (repeatable) ║ ║ --parallel, -p Run
Explore related resources
Frequently asked questions
What is waza?
waza is a open-source AI agent skill with Copy skill directory. WORKFLOW SKILL - Evaluate AI agent skills using structured benchmarks with YAML specs, fixture isolation, and pluggable validators.
Who is waza best for?
waza is best for reusing agent instructions, scripts, and references, testing workflows.
How do I install waza?
Install or run waza using Copy skill directory. Check waza for the latest setup command.
Is waza actively maintained?
waza may need a closer maintenance check before production use.
Auto-fetched from GitHub.