eval-analysis Skill
Comprehensive guide for analyzing SABER evaluation results — model comparison, agent architecture comparison, domain-specific analysis, and cross-domain aggregate analysis. Use this when asked to analyze eval results, compare models, generate visualizations, interpret scores, investigate cost-efficiency trade-offs, or Published by microsoft in ACESEvals.
What is eval-analysis Skill?
Comprehensive guide for analyzing SABER evaluation results — model comparison, agent architecture comparison, domain-specific analysis, and cross-domain aggregate analysis. Use this when asked to analyze eval results, compare models, generate visualizations, interpret scores, investigate cost-efficiency trade-offs, or Published by microsoft in ACESEvals. This profile combines repository metadata with install, compatibility, and usage signals so developers can quickly decide whether it fits their agent workflow before opening the source repository.
Automated repository signals based on public metadata such as recency, license, installation evidence, and adoption. These are not a security audit or endorsement.
Key capabilities
- Includes SKILL.md support
- Reusable instructions support
- Data analysis
- Data analysis use cases
Technical details
- Install or run with Copy skill directory
When to use eval-analysis Skill
- Use it for data analysis.
Built with
Editorial notes
Source
- Creator: microsoft
- Repository: microsoft/ACESEvals
- Skill file: .github/skills/eval-analysis/SKILL.md
What it does
Comprehensive guide for analyzing SABER evaluation results — model comparison, agent architecture comparison, domain-specific analysis, and cross-domain aggregate analysis. Use this when asked to analyze eval results, compare models, generate visualizations, interpret scores, investigate cost-efficiency trade-offs, or
Skill instructions
Eval Analysis Skill Overview SABER evaluation analysis is a structured data science framework for comparing AI agent performance across cybersecurity benchmark domains. The analysis pipeline transforms .eval log files into actionable insights through 12+ standardized experiments plus domain-specific analyses. Getting Started: Sample Eval Data No evals to analyze yet? The repo ships with pre-run sample .eval files and can auto-download more from HuggingFace. Option 1: Use bundled sample evals (fastest) The evalsamples/ directory contains pre-run evaluation results for all 3 domains across 5 models and 3 agent architectures: evalsamples/ ├── excytin/ 20 eval files (5 models × default + reasoning baselines + 3 agents) ├── cybench/ 17 eval files └── ctirealm/ 11 eval files Models included: Claude Haiku 4.5, Claude Sonnet 4.6, Claude Opus 4.6, GPT-5.4, GPT-5.4-mini Agent architectures included: React, GH Copilot, Claude Code (for Sonnet 4.6) Baselines: No-reasoning/no-thinking variants for
Explore related resources
Frequently asked questions
What is eval-analysis?
eval-analysis is a open-source AI agent skill with Copy skill directory. Comprehensive guide for analyzing SABER evaluation results — model comparison, agent architecture comparison, domain-specific analysis, and cross-domain aggregate analysis.
Who is eval-analysis best for?
eval-analysis is best for reusing agent instructions, scripts, and references, data analysis workflows.
How do I install eval-analysis?
Install or run eval-analysis using Copy skill directory. Check eval-analysis for the latest setup command.
Is eval-analysis actively maintained?
eval-analysis may need a closer maintenance check before production use.
Auto-fetched from GitHub.