Favicon of ai-eval-failure-analysis

ai-eval-failure-analysis Skill

AI Agent SkillALOpen source

Analyze Business Central AI Test Toolkit (AIT) evaluation results to explain WHY an eval run failed. Use when the user has an AIT run (a version/suite, or a saved aitTestLogEntries JSON export) and wants failure analysis, a confusion matrix, precision/recall/F1, specificity, MatchRate, accuracy, sibling-confusion, mode Published by microsoft in BCTech.

What is ai-eval-failure-analysis Skill?

Analyze Business Central AI Test Toolkit (AIT) evaluation results to explain WHY an eval run failed. Use when the user has an AIT run (a version/suite, or a saved aitTestLogEntries JSON export) and wants failure analysis, a confusion matrix, precision/recall/F1, specificity, MatchRate, accuracy, sibling-confusion, mode Published by microsoft in BCTech. This profile combines repository metadata with install, compatibility, and usage signals so developers can quickly decide whether it fits their agent workflow before opening the source repository.

Trust signal
95/100
Maintenance signal
90/100
Adoption signal
71/100

Automated repository signals based on public metadata such as recency, license, installation evidence, and adoption. These are not a security audit or endorsement.

Key capabilities

  • Includes SKILL.md support
  • Reusable instructions support
  • Testing
  • Data analysis
  • Testing use cases
  • Data analysis use cases

Technical details

Copy skill directory
  • Install or run with Copy skill directory

When to use ai-eval-failure-analysis Skill

  • Use it for testing.
  • Use it for data analysis.

Built with

ALCopy skill directory

Editorial notes

Source

  • Creator: microsoft
  • Repository: microsoft/BCTech
  • Skill file: samples/AIFailureAnalysis/SKILL.md

What it does

Analyze Business Central AI Test Toolkit (AIT) evaluation results to explain WHY an eval run failed. Use when the user has an AIT run (a version/suite, or a saved aitTestLogEntries JSON export) and wants failure analysis, a confusion matrix, precision/recall/F1, specificity, MatchRate, accuracy, sibling-confusion, mode

Skill instructions

AI Eval Failure Analysis (Business Central AIT) You help developers understand why a Business Central AI Test Toolkit (AIT) run failed and how strong a model actually is on the task — going well beyond the toolkit's built-in pass/fail count. When to use this skill Use it whenever the user references an AIT run and wants more than "X/Y passed": failure root-causing, richer metrics, or a comparison across model/prompt versions. Inputs the skill understands: - A saved export of the AIT log (aitTestLogEntries rows as JSON) — preferred, works anywhere. - A live run identified by suite code + codeunit + version, fetched from the AIT Toolkit OData API (only on a box with the server + credentials). The AIT log shape (what you are reasoning over) Each row in the log is one dataset line and carries: | Field | Meaning | |-------|---------| | inputData | JSON string: { "input": <prompt context, "expectedoutput": <ground truth } | | outputData | JSON string: { "answer": <model output } | | status |

Explore related resources

Frequently asked questions

What is ai-eval-failure-analysis?

ai-eval-failure-analysis is a open-source AI agent skill with Copy skill directory. Analyze Business Central AI Test Toolkit (AIT) evaluation results to explain WHY an eval run failed.

Who is ai-eval-failure-analysis best for?

ai-eval-failure-analysis is best for reusing agent instructions, scripts, and references, testing workflows, data analysis workflows.

How do I install ai-eval-failure-analysis?

Install or run ai-eval-failure-analysis using Copy skill directory. Check ai-eval-failure-analysis for the latest setup command.

Is ai-eval-failure-analysis actively maintained?

ai-eval-failure-analysis may need a closer maintenance check before production use.

Share:

Stars
735
Forks
383
Last commit
1 month ago
Repository age
8 years
License
MIT

Auto-fetched from GitHub.

Ad
Favicon

 

  
 

Similar to ai-eval-failure-analysis