waza Skill
WORKFLOW SKILL - Evaluate AI agent skills using structured benchmarks with YAML specs, fixture isolation, and pluggable validators. USE FOR: run waza, waza help, run eval, run benchmark, evaluate skill, test agent, generate eval suite, init eval, compare results, score agent, agent evaluation, skill testing, cross-mode Published by microsoft in waza.
Decision snapshot
Is this a fit?
Testing, Includes SKILL.md, Reusable instructions
Compatibility not yet detected.
Permission behavior not yet detected.
Copy skill directory
2 months ago · MIT license
No specific cautions were detected. Review the source and requested permissions before installing.
What is waza Skill?
WORKFLOW SKILL - Evaluate AI agent skills using structured benchmarks with YAML specs, fixture isolation, and pluggable validators. USE FOR: run waza, waza help, run eval, run benchmark, evaluate skill, test agent, generate eval suite, init eval, compare results, score agent, agent evaluation, skill testing, cross-mode Published by microsoft in waza. This profile combines repository metadata with install, compatibility, and usage signals so developers can quickly decide whether it fits their agent workflow before opening the source repository.
Automated repository signals based on public metadata such as recency, license, installation evidence, and adoption. These are not a security audit or endorsement. See how SkillIndex evaluates profiles.
Key capabilities
- Includes SKILL.md support
- Reusable instructions support
- Testing
- Testing use cases
Declared skill metadata
- Source file: skills/waza/SKILL.md
These fields retain source and confidence evidence from the indexed SKILL.md.
Compatibility and setup
- Install or run with Copy skill directory
When to use waza Skill
- Use it for testing.
Built with
Editorial notes
Source
- Creator: microsoft
- Repository: microsoft/waza
- Skill file: skills/waza/SKILL.md
What it does
WORKFLOW SKILL - Evaluate AI agent skills using structured benchmarks with YAML specs, fixture isolation, and pluggable validators. USE FOR: run waza, waza help, run eval, run benchmark, evaluate skill, test agent, generate eval suite, init eval, compare results, score agent, agent evaluation, skill testing, cross-mode
Skill instructions
Waza "The way of technique — measure, refine, master." A Go CLI tool for evaluating AI agent skills through structured benchmarks. Define test cases in YAML, run them against agent engines, and validate results with pluggable scoring validators. Help When user says "waza help" or asks how to use waza: ╔══════════════════════════════════════════════════════════════════╗ ║ WAZA - CLI Tool for Evaluating Agent Skills ║ ╠══════════════════════════════════════════════════════════════════╣ ║ ║ ║ COMMANDS: ║ ║ waza run <eval.yaml Run an evaluation benchmark ║ ║ waza init [directory] Initialize a new eval suite ║ ║ waza generate <SKILL.md Generate eval from SKILL.md ║ ║ waza compare <r1 <r2 ... Compare result files ║ ║ waza dev [skill-path] Improve SKILL.md compliance ║ ║ ║ ║ RUN FLAGS: ║ ║ --context-dir, -c Fixtures directory (default: ./fixtures) ║ ║ --output, -o Save results JSON to file ║ ║ --verbose, -v Verbose output ║ ║ --task, -t Filter tasks by name (repeatable) ║ ║ --parallel, -p Run
Verified compatibility and discovery
Frequently asked questions
What is waza?
waza is a open-source AI agent skill with Copy skill directory. WORKFLOW SKILL - Evaluate AI agent skills using structured benchmarks with YAML specs, fixture isolation, and pluggable validators.
Who is waza best for?
waza is best for reusing agent instructions, scripts, and references, testing workflows.
How do I install waza?
Install or run waza using Copy skill directory. Check waza for the latest setup command.
Is waza actively maintained?
waza may need a closer maintenance check before production use.
Project health auto-fetched from the source repository.
Maintain this resource?
Review this source-backed profile, send a correction with evidence, or link to it from your documentation. Claims verify your relationship to the project; profile facts still require source evidence and editorial review.