eval-analysis Skill
Comprehensive guide for analyzing SABER evaluation results — model comparison, agent architecture comparison, domain-specific analysis, and cross-domain aggregate analysis. Use this when asked to analyze eval results, compare models, generate visualizations, interpret scores, investigate cost-efficiency trade-offs, or Published by microsoft in ACESEvals.
Decision snapshot
Is this a fit?
Data analysis, Includes SKILL.md, Reusable instructions
Compatibility not yet detected.
Permission behavior not yet detected.
Copy skill directory
2 months ago · MIT license
No specific cautions were detected. Review the source and requested permissions before installing.
What is eval-analysis Skill?
Comprehensive guide for analyzing SABER evaluation results — model comparison, agent architecture comparison, domain-specific analysis, and cross-domain aggregate analysis. Use this when asked to analyze eval results, compare models, generate visualizations, interpret scores, investigate cost-efficiency trade-offs, or Published by microsoft in ACESEvals. This profile combines repository metadata with install, compatibility, and usage signals so developers can quickly decide whether it fits their agent workflow before opening the source repository.
Automated repository signals based on public metadata such as recency, license, installation evidence, and adoption. These are not a security audit or endorsement. See how SkillIndex evaluates profiles.
Key capabilities
- Includes SKILL.md support
- Reusable instructions support
- Data analysis
- Data analysis use cases
Declared skill metadata
- Source file: .github/skills/eval-analysis/SKILL.md
These fields retain source and confidence evidence from the indexed SKILL.md.
Compatibility and setup
- Install or run with Copy skill directory
When to use eval-analysis Skill
- Use it for data analysis.
Built with
Editorial notes
Source
- Creator: microsoft
- Repository: microsoft/ACESEvals
- Skill file: .github/skills/eval-analysis/SKILL.md
What it does
Comprehensive guide for analyzing SABER evaluation results — model comparison, agent architecture comparison, domain-specific analysis, and cross-domain aggregate analysis. Use this when asked to analyze eval results, compare models, generate visualizations, interpret scores, investigate cost-efficiency trade-offs, or
Skill instructions
Eval Analysis Skill Overview SABER evaluation analysis is a structured data science framework for comparing AI agent performance across cybersecurity benchmark domains. The analysis pipeline transforms .eval log files into actionable insights through 12+ standardized experiments plus domain-specific analyses. Getting Started: Sample Eval Data No evals to analyze yet? The repo ships with pre-run sample .eval files and can auto-download more from HuggingFace. Option 1: Use bundled sample evals (fastest) The evalsamples/ directory contains pre-run evaluation results for all 3 domains across 5 models and 3 agent architectures: evalsamples/ ├── excytin/ 20 eval files (5 models × default + reasoning baselines + 3 agents) ├── cybench/ 17 eval files └── ctirealm/ 11 eval files Models included: Claude Haiku 4.5, Claude Sonnet 4.6, Claude Opus 4.6, GPT-5.4, GPT-5.4-mini Agent architectures included: React, GH Copilot, Claude Code (for Sonnet 4.6) Baselines: No-reasoning/no-thinking variants for
Verified compatibility and discovery
Frequently asked questions
What is eval-analysis?
eval-analysis is a open-source AI agent skill with Copy skill directory. Comprehensive guide for analyzing SABER evaluation results — model comparison, agent architecture comparison, domain-specific analysis, and cross-domain aggregate analysis.
Who is eval-analysis best for?
eval-analysis is best for reusing agent instructions, scripts, and references, data analysis workflows.
How do I install eval-analysis?
Install or run eval-analysis using Copy skill directory. Check eval-analysis for the latest setup command.
Is eval-analysis actively maintained?
eval-analysis may need a closer maintenance check before production use.
Project health auto-fetched from the source repository.
Maintain this resource?
Review this source-backed profile, send a correction with evidence, or link to it from your documentation. Claims verify your relationship to the project; profile facts still require source evidence and editorial review.