How does AWS want teams to test skill-equipped AI agents?

AWS is adding skill-focused evaluators for agents. What do they measure, and what remains unproven?

How does AWS want teams to test skill-equipped AI agents?
Editorial illustration of AWS AgentCore and Strands Evals agent testing. Illustration generated by Eazzy Tech News. Not a photograph of a real event.

AWS says Strands Evals and Amazon Bedrock AgentCore Evaluations can now test whether an AI agent chose the right domain skill and whether it followed that skill's instructions. The useful part is the narrower claim: this is an evaluation method for agent procedure failures, not a guarantee that a deployed agent is safe or correct.

The September 22 AWS Machine Learning Blog post frames skills as reusable instruction packages for tasks such as redacting contracts, reconciling invoices or following engineering conventions. AWS argues that general output-quality checks can miss two practical failures: an agent may load the wrong skill, or it may load the right skill and skip required steps.

AWS-published diagram showing Skill Selection Accuracy scoring an agent's selected skill.
AWS illustrates Skill Selection Accuracy as a check on whether an agent selected the appropriate skill for a user task. Source: AWS.

What AWS is adding

The announcement centers on three evaluators. Skill Selection Accuracy checks whether each invoked skill was the appropriate one for the task. Skill Instruction Following rates how completely the agent followed the prescribed steps for an invoked skill. A third Strands-only check, Skill Invoked, deterministically verifies whether a named skill loaded successfully.

That separation matters because the fixes are different. If a router chooses the wrong skill, teams may need clearer skill descriptions or less overlap in the catalog. If the right skill is selected but steps are skipped, the issue may be the skill instructions, the harness, tool failures, context handling or the model used by the agent.

Evaluator Where AWS says it runs Question it answers
Skill Selection Accuracy Strands Evals and AgentCore Evaluations Did the agent pick a suitable skill for the task?
Skill Instruction Following Strands Evals and AgentCore Evaluations Did the agent follow the steps inside the selected skill?
Skill Invoked Strands Evals Was a named skill loaded in the test run?

Why this is a real agent story

Agent reliability is often discussed as if the final answer is the only artifact worth judging. AWS is pointing at a more operational layer: the trace of what the agent loaded, what it tried, and which prescribed steps it completed. The source says Strands Evals can work from recorded trajectories during development, while AgentCore Evaluations can use OpenTelemetry traces from agents hosted on AgentCore runtime or elsewhere.

AWS also lists several harness signals that skill extraction can recognize at launch, including SKILL.md file reads and integrations from frameworks such as Strands Agents, LangGraph Deep Agents, Google ADK, Claude Agent SDK, OpenAI Agents SDK, Codex, Gemini CLI and OpenHands. That makes the announcement relevant beyond a single AWS sample app, although teams still need to instrument and calibrate their own cases.

AWS-published diagram showing Skill Instruction Following labels from fully followed to not followed.
AWS depicts Skill Instruction Following as a step-level rating for how completely an agent followed an invoked skill. Source: AWS.

What is not settled

The post is a technical how-to and announcement, not an independent benchmark across agent frameworks. Its examples use an HR assistant and developer commands to show how teams can define cases, run evaluations and gate deployments. The important qualification is that judge-based evaluators still need good test cases, appropriate thresholds and review of false positives or false negatives.

For AI teams, the practical takeaway is concrete: if agents are going to depend on modular skills, evaluation needs to test the route and the execution path separately. AWS is adding tooling for that workflow; the evidence of reliability will still come from how each team records traces, chooses cases, handles failures and decides what score is good enough to ship.

Sources