Favicon of eval-analysis

eval-analysis Skill

AI Agent SkillJupyter NotebookOpen source

Comprehensive guide for analyzing SABER evaluation results — model comparison, agent architecture comparison, domain-specific analysis, and cross-domain aggregate analysis. Use this when asked to analyze eval results, compare models, generate visualizations, interpret scores, investigate cost-efficiency trade-offs, or Published by microsoft in ACESEvals.

What is eval-analysis Skill?

Comprehensive guide for analyzing SABER evaluation results — model comparison, agent architecture comparison, domain-specific analysis, and cross-domain aggregate analysis. Use this when asked to analyze eval results, compare models, generate visualizations, interpret scores, investigate cost-efficiency trade-offs, or Published by microsoft in ACESEvals. This profile combines repository metadata with install, compatibility, and usage signals so developers can quickly decide whether it fits their agent workflow before opening the source repository.

Trust signal
95/100
Maintenance signal
90/100
Adoption signal
17/100

Automated repository signals based on public metadata such as recency, license, installation evidence, and adoption. These are not a security audit or endorsement.

Key capabilities

  • Includes SKILL.md support
  • Reusable instructions support
  • Data analysis
  • Data analysis use cases

Technical details

Copy skill directory
  • Install or run with Copy skill directory

When to use eval-analysis Skill

  • Use it for data analysis.

Built with

Jupyter NotebookCopy skill directory

Editorial notes

Source

  • Creator: microsoft
  • Repository: microsoft/ACESEvals
  • Skill file: .github/skills/eval-analysis/SKILL.md

What it does

Comprehensive guide for analyzing SABER evaluation results — model comparison, agent architecture comparison, domain-specific analysis, and cross-domain aggregate analysis. Use this when asked to analyze eval results, compare models, generate visualizations, interpret scores, investigate cost-efficiency trade-offs, or

Skill instructions

Eval Analysis Skill Overview SABER evaluation analysis is a structured data science framework for comparing AI agent performance across cybersecurity benchmark domains. The analysis pipeline transforms .eval log files into actionable insights through 12+ standardized experiments plus domain-specific analyses. Getting Started: Sample Eval Data No evals to analyze yet? The repo ships with pre-run sample .eval files and can auto-download more from HuggingFace. Option 1: Use bundled sample evals (fastest) The evalsamples/ directory contains pre-run evaluation results for all 3 domains across 5 models and 3 agent architectures: evalsamples/ ├── excytin/ 20 eval files (5 models × default + reasoning baselines + 3 agents) ├── cybench/ 17 eval files └── ctirealm/ 11 eval files Models included: Claude Haiku 4.5, Claude Sonnet 4.6, Claude Opus 4.6, GPT-5.4, GPT-5.4-mini Agent architectures included: React, GH Copilot, Claude Code (for Sonnet 4.6) Baselines: No-reasoning/no-thinking variants for

Explore related resources

Frequently asked questions

What is eval-analysis?

eval-analysis is a open-source AI agent skill with Copy skill directory. Comprehensive guide for analyzing SABER evaluation results — model comparison, agent architecture comparison, domain-specific analysis, and cross-domain aggregate analysis.

Who is eval-analysis best for?

eval-analysis is best for reusing agent instructions, scripts, and references, data analysis workflows.

How do I install eval-analysis?

Install or run eval-analysis using Copy skill directory. Check eval-analysis for the latest setup command.

Is eval-analysis actively maintained?

eval-analysis may need a closer maintenance check before production use.

Share:

Stars
4
Forks
1
Last commit
12 days ago
Repository age
4 months
License
MIT

Auto-fetched from GitHub.

Ad
Favicon

 

  
 

Similar to eval-analysis