How does AWS want teams to test skill-equipped AI agents?
AWS is adding skill-focused evaluators for agents. What do they measure, and what remains unproven?
AWS says Strands Evals and Amazon Bedrock AgentCore Evaluations can now test whether an AI agent chose the right domain skill and whether it followed that skill's instructions. The useful part is the narrower claim: this is an evaluation method for agent procedure failures, not a guarantee that a deployed agent is safe or correct.
The September 22 AWS Machine Learning Blog post frames skills as reusable instruction packages for tasks such as redacting contracts, reconciling invoices or following engineering conventions. AWS argues that general output-quality checks can miss two practical failures: an agent may load the wrong skill, or it may load the right skill and skip required steps.

What AWS is adding
The announcement centers on three evaluators. Skill Selection Accuracy checks whether each invoked skill was the appropriate one for the task. Skill Instruction Following rates how completely the agent followed the prescribed steps for an invoked skill. A third Strands-only check, Skill Invoked, deterministically verifies whether a named skill loaded successfully.
That separation matters because the fixes are different. If a router chooses the wrong skill, teams may need clearer skill descriptions or less overlap in the catalog. If the right skill is selected but steps are skipped, the issue may be the skill instructions, the harness, tool failures, context handling or the model used by the agent.
| Evaluator | Where AWS says it runs | Question it answers |
|---|---|---|
| Skill Selection Accuracy | Strands Evals and AgentCore Evaluations | Did the agent pick a suitable skill for the task? |
| Skill Instruction Following | Strands Evals and AgentCore Evaluations | Did the agent follow the steps inside the selected skill? |
| Skill Invoked | Strands Evals | Was a named skill loaded in the test run? |
Why this is a real agent story
Agent reliability is often discussed as if the final answer is the only artifact worth judging. AWS is pointing at a more operational layer: the trace of what the agent loaded, what it tried, and which prescribed steps it completed. The source says Strands Evals can work from recorded trajectories during development, while AgentCore Evaluations can use OpenTelemetry traces from agents hosted on AgentCore runtime or elsewhere.
AWS also lists several harness signals that skill extraction can recognize at launch, including SKILL.md file reads and integrations from frameworks such as Strands Agents, LangGraph Deep Agents, Google ADK, Claude Agent SDK, OpenAI Agents SDK, Codex, Gemini CLI and OpenHands. That makes the announcement relevant beyond a single AWS sample app, although teams still need to instrument and calibrate their own cases.

What is not settled
The post is a technical how-to and announcement, not an independent benchmark across agent frameworks. Its examples use an HR assistant and developer commands to show how teams can define cases, run evaluations and gate deployments. The important qualification is that judge-based evaluators still need good test cases, appropriate thresholds and review of false positives or false negatives.
For AI teams, the practical takeaway is concrete: if agents are going to depend on modular skills, evaluation needs to test the route and the execution path separately. AWS is adding tooling for that workflow; the evidence of reliability will still come from how each team records traces, chooses cases, handles failures and decides what score is good enough to ship.
Comments ()