Eval Audit And Sweep Skill
<!-- Copyright 2026 Anthropic PBC -- <!-- SPDX-License-Identifier: Apache-2.0 -- Published by anthropics in cwc-workshops.
Decision snapshot
Is this a fit?
Testing, Security review, Includes SKILL.md, Reusable instructions
Compatibility not yet detected.
Permission behavior not yet detected.
Copy skill directory
19 days ago · Apache-2.0 license
No specific cautions were detected. Review the source and requested permissions before installing.
What is Eval Audit And Sweep Skill?
<!-- Copyright 2026 Anthropic PBC -- <!-- SPDX-License-Identifier: Apache-2.0 -- Published by anthropics in cwc-workshops. This profile combines repository metadata with install, compatibility, and usage signals so developers can quickly decide whether it fits their agent workflow before opening the source repository.
Automated repository signals based on public metadata such as recency, license, installation evidence, and adoption. These are not a security audit or endorsement. See how SkillIndex evaluates profiles.
Key capabilities
- Includes SKILL.md support
- Reusable instructions support
- Testing
- Security review
- Testing use cases
- Security review use cases
Declared skill metadata
- Source file: rightmodel/.claude/skills/eval-audit-and-sweep/SKILL.md
These fields retain source and confidence evidence from the indexed SKILL.md.
Compatibility and setup
- Install or run with Copy skill directory
When to use Eval Audit And Sweep Skill
- Use it for testing.
- Use it for security review.
Built with
Editorial notes
Source
- Creator: anthropics
- Repository: anthropics/cwc-workshops
- Skill file: rightmodel/.claude/skills/eval-audit-and-sweep/SKILL.md
What it does
<!-- Copyright 2026 Anthropic PBC -- <!-- SPDX-License-Identifier: Apache-2.0 -- ## Skill instructions <!-- Copyright 2026 Anthropic PBC -- <!-- SPDX-License-Identifier: Apache-2.0 -- --- name: eval-audit-and-sweep description: This skill should be used when a user wants to (a) audit the quality and reliability of an existing LLM evaluation suite, or (b) determine which Claude model and inference parameters give the best quality-per-dollar and quality-per-second for their specific task by running a parameter sweep over that eval. Applicable to any eval framework (custom harnesses, tau-bench, inspect-ai, pytest-based, etc.) since the guidance is framework-agnostic. --- Eval Audit and Sweep This skill is an example exercise for the "Picking the Right Model" workshop during Code with Claude. It is a two-phase playbook for getting trustworthy cost-quality numbers out of an existing LLM eval. Phase 1 audits the eval for common reliability issues. Phase 2 wraps it in a model/parameter sweep and produces a recommendation. The phases are independent: a user may ask for only the audit, only theVerified compatibility and discovery
Frequently asked questions
What is Eval Audit And Sweep?
Eval Audit And Sweep is a open-source AI agent skill with Copy skill directory. <!-- Copyright 2026 Anthropic PBC -- <!-- SPDX-License-Identifier: Apache-2.0 --
Who is Eval Audit And Sweep best for?
Eval Audit And Sweep is best for reusing agent instructions, scripts, and references, testing workflows, security review workflows.
How do I install Eval Audit And Sweep?
Install or run Eval Audit And Sweep using Copy skill directory. Check Eval Audit And Sweep for the latest setup command.
Is Eval Audit And Sweep actively maintained?
Eval Audit And Sweep may need a closer maintenance check before production use.
Project health auto-fetched from the source repository.
Maintain this resource?
Review this source-backed profile, send a correction with evidence, or link to it from your documentation. Claims verify your relationship to the project; profile facts still require source evidence and editorial review.