Build, deploy and operate agents with AWS AWS Startups x JAM · Tue 20 Oct 2026 · AWS Singapore

11:35–12:25

3. Intake: is your agent any good?

The misread. A generic judge checks the answer against the tools. Yours checks it against the invoice.

The brief

Intake is the invoice inbox of a small Singapore company. A scanner reads each supplier invoice, picks out who it's from and how much it's for, and puts it in a queue for the finance team to approve. Finance ask an agent about the queue: what's waiting, which ones are big, does anything look odd.

Last week the scanner misread an invoice from Lucky Star Print. The paper said S$1,280.00. The scanner missed the decimal point and recorded S$128,000.00. The agent read that number back to finance as if it were fine.

The scanner will be fixed. Your job is to catch the agent when it gets an amount wrong, before finance does, and make it more careful in the meantime.

What you'll learn

An agent that sounds confident isn't the same as an agent that's right. By the end of this scenario you'll be able to:

  1. Score an agent's answers with AgentCore Evaluations, using built-in judges for correctness, helpfulness and tool choice.
  2. Explain why a generic judge can pass an answer that's wrong for your business.
  3. Write a custom evaluator that checks your own rule, here that every amount matches what's printed on the invoice.
  4. Make the agent more careful with config alone: a stricter prompt and a sandboxed Code Interpreter for arithmetic, as a new harness version you can roll back.
  5. Score live traffic continuously with online evaluation, and pick a sensible sampling rate for production.

AgentCore pieces: Harness, Gateway, Evaluations, Code Interpreter

You'll leave with: a way to find out your agent is wrong before your users do, and a before-and-after score to prove a fix worked.

Time: 50 minutes. Before you start: Kopi Run, or at least its harness steps, because you'll build this agent the same way.

Every step here works on both paths, so use whichever of the Console or Terminal tabs suits you. Stay on the same path for the whole scenario: the Terminal steps build on the IntakeQA project that the first step creates, so they won't work on an agent you built in the console.

Deploy the finance agent

You've done this once, so it should feel familiar. The first prompt is short and trusting:

You answer finance questions about supplier invoices in the approval queue.
Use list_queue and get_invoice to look things up. Be concise.

The tools come from a Lambda function called agentcore-workshop-intake-api. get_invoice returns three things worth knowing about: extracted, which is what the scanner read; raw_total_text, the total exactly as printed on the invoice; and confidence, how sure the scanner was, from 0 to 1.

Start by copying the Lambda ARN. Run this in the code editor terminal and copy what it prints:

echo $INTAKE_API_ARN
  1. Choose Gateways, then Create Gateway. Step 1: replace the generated name with intake-gw and keep Create default role. Choose Next.

    Screenshot: Define gateway details

    Create gateway, step 1: gateway name intake-gw and Create default role selected

  2. Step 2: change Inbound Auth type to Use IAM permissions. Choose Next.

  3. Step 3: leave MCP target as it is. Target name: intake-api. Target type: Lambda ARN, and paste the ARN you copied. Target schema: Define an inline schema, then paste the tool list below. Leave Outbound Auth as IAM Role. Choose Next.

    Screenshot: Add targets, with the Intake tool list pasted

    Create gateway, step 3: target intake-api, Lambda ARN, inline schema editor with the Intake tools

  4. Step 4: review and create the gateway, then wait for it to show Ready.

  5. Choose Harness, open the arrow on Quick create Harness, and choose Advanced create Harness. Name: intake_agent. Keep Bedrock and Claude Sonnet 4.6, and replace the default system prompt with the prompt above.

    Screenshot: name, model and the first system prompt

    Create Harness: name intake_agent, Claude Sonnet 4.6, Intake system prompt

  6. Leave managed memory on. Under Tools, switch on Gateway and select intake-gw.

    Screenshot: memory and tools

    Create Harness: managed memory enabled, Gateway tool switched on

  7. Under Advanced configurations, Invocation limits, change Max Iterations to 20. Leave Create default role under Permissions.

    Screenshot: invocation limits

    Create Harness: Invocation limits, Max Iterations 20

  8. Choose Create Harness and wait for Ready. It takes three to five minutes while memory is provisioned.

Tool list to paste
[
  {
    "name": "list_queue",
    "description": "List invoices waiting for approval: invoice ID, supplier, currency and extracted total.",
    "inputSchema": { "type": "object", "properties": {} }
  },
  {
    "name": "get_invoice",
    "description": "Get one invoice: the extracted fields, the total exactly as printed on the document (raw_total_text), and the extractor's confidence from 0 to 1.",
    "inputSchema": {
      "type": "object",
      "properties": { "invoice_id": { "type": "string" } },
      "required": ["invoice_id"]
    }
  }
]
cd $WORKSHOP &&
agentcore create --project-name IntakeQA --no-agent &&
cd IntakeQA &&
agentcore add gateway --name intake-gw --authorizer-type AWS_IAM &&
agentcore add gateway-target \
  --type lambda-function-arn \
  --name intake-api \
  --lambda-arn "$INTAKE_API_ARN" \
  --tool-schema-file "$WORKSHOP/intake/tools.json" \
  --gateway intake-gw &&
agentcore deploy -y &&
agentcore add harness \
  --name intake_agent \
  --model-provider bedrock \
  --system-prompt "$(cat $WORKSHOP/intake/system-prompt-v1.md)" \
  --max-iterations 20 &&
agentcore add tool --harness intake_agent --type agentcore_gateway --name intake --gateway intake-gw &&
agentcore deploy -y

While it deploys, have a look at the files:

cat $WORKSHOP/intake/tools.json $WORKSHOP/intake/system-prompt-v1.md

Ask it some questions

Open Harness playground (under Test), select intake_agent and choose Test Harness. Ask each of these in its own new session, so each one is scored separately later:

Screenshot: Harness playground

Harness playground: Select a Harness and Test Harness

  1. What's the total value of invoices waiting for approval?
  2. Which invoices are over S$5,000?
  3. How much is the Lucky Star Print invoice?
  4. Is anything in the queue unusual?
cat $WORKSHOP/intake/questions.txt &&
$WORKSHOP/scripts/intake-ask.sh intake_agent $WORKSHOP/intake/questions.txt

The script asks each question in its own session and prints the answers.

Read the answers, especially the one about Lucky Star Print.

These sessions need five to ten minutes before Evaluations can score them. The next step fills that time, so carry straight on.

Write an evaluator that knows your business

What is AgentCore Evaluations, and why use it?

Ordinary tests check that code returns the right value. An agent's answers vary from run to run, so you need something that judges quality instead: was it correct, was it helpful, did it pick the right tool. Evaluations scores your agent's traces with a model acting as the judge (LLM-as-a-judge), using built-in evaluators for general qualities and custom ones you write for your own rules. You can score on demand, in a batch, or continuously on a sample of live traffic.

That's how you catch a regression after a prompt change before your users do, and compare versions with numbers rather than impressions.

Sources: AgentCore Evaluations

AgentCore Evaluations uses a model as a judge to score whole sessions, single turns or single tool calls. It comes with built-in evaluators for general qualities such as correctness, helpfulness and whether the agent picked the right tool. You'll use those in a moment. First, write one that checks the thing finance actually cares about:

You are reviewing an answer from a finance assistant about supplier invoices.

Conversation and tool results: {context}
Answer being reviewed: {assistant_turn}

Pass the answer only if both of these are true:
1. Every invoice amount the assistant states matches the amount printed on
   the invoice (raw_total_text in the tool results), not just the scanned
   figure (extracted.total).
2. Any invoice with a scan confidence below 0.9 is flagged for human review.

Fail it if either is false. If the answer states no invoice amounts, pass it.

{context} and {assistant_turn} are filled in by Evaluations for each turn it scores.

  1. Choose Evaluations in the navigation pane, open the Custom evaluators tab, and choose Create custom evaluator.

  2. Evaluator name: replace the generated name with AmountCheck. Leave the definition type as Custom prompt.

  3. The Instruction box is pre-filled with the Faithfulness template. Choose Clear, then paste the text above.

  4. Model: search for Sonnet 4.6 and choose Global Anthropic Claude Sonnet 4.6.

    Screenshot: evaluator name, prompt and model

    Create custom evaluator: name AmountCheck, Custom prompt, Instruction box and Model picker

  5. Scale type: choose Define scale as string values. Delete rows until two are left, then set them to Pass (Every amount matches the document, and low-confidence invoices are flagged for review.) and Fail (An amount is misread, or a low-confidence invoice is not flagged.).

  6. Evaluation level: leave Trace selected.

    Screenshot: scale and evaluation level

    Scale definitions Pass and Fail, evaluation level Trace

  7. Choose Create custom evaluator.

This adds the evaluator to the IntakeQA project you created in "Deploy the finance agent". Check it's there first: ls $WORKSHOP/IntakeQA should list agentcore and app.

cd $WORKSHOP/IntakeQA &&
agentcore add evaluator \
  --name AmountCheck \
  --level TRACE \
  --rating-scale pass-fail \
  --model global.anthropic.claude-sonnet-4-6 \
  --instructions "$(cat $WORKSHOP/intake/evaluators/amount-check.txt)" &&
agentcore deploy -y

Score the answers

Now run the built-in evaluators and your own together against the sessions you created earlier.

In the console, a one-off evaluation over existing traces is called a batch evaluation.

  1. In Evaluations, open the Batch evaluation tab and choose Create batch evaluation.

  2. Batch evaluation name: intake_first_run.

  3. Under the built-in evaluators, tick Correctness and Helpfulness (response quality) and Tool selection accuracy (component level). Under Custom evaluators, tick AmountCheck.

    Screenshot: choosing evaluators

    Create batch evaluation with Correctness, Helpfulness and Tool selection accuracy ticked

  4. Input data source: leave Define with a runtime agent endpoint and choose intake_agent under Choose agent.

    Screenshot: input data source

    Input data source: Define with a runtime agent endpoint, Choose agent

  5. Choose Create batch evaluation, then open it from the list to see the scores.

$WORKSHOP/scripts/eval-harness.sh intake_agent \
  Builtin.Correctness Builtin.Helpfulness Builtin.ToolSelectionAccuracy AmountCheck

agentcore run eval scores a Runtime, and a harness isn't one in the project's eyes (you'd see No runtimes defined in agentcore.json). Every harness runs on an AgentCore Runtime underneath, though, so the script looks up that Runtime and your evaluator's ARN, and runs the CLI's standalone mode: agentcore run eval --runtime-arn ... --evaluator-arn .... It runs one evaluator at a time, so if one can't finish yet, the others still give you scores.

If no sessions are found, the traces aren't ready yet. Give it two more minutes and try again.

Compare the scores for the Lucky Star Print answer. The built-in correctness judge probably passed it: it checks the answer against what the tools returned, and the tools said S$128,000, so repeating that counts as correct. A generic judge can't know your scanner is wrong. AmountCheck does, and it should fail that answer along with any answer that added the bad figure into a total.

Make the agent more careful

What is Code Interpreter, and why use it?

Language models predict text. They aren't calculators, and adding up a list of invoice totals in their head is exactly where mistakes creep in. Code Interpreter gives the agent a sandbox where it can write and run code (Python, JavaScript or TypeScript) and use the result. The sandbox is isolated, so running code the model wrote doesn't put your systems at risk, and it's built in: you switch it on as a tool, with nothing to host.

Sources: AgentCore Code Interpreter

The scanner fix belongs to engineering. You can still make the agent safer today with two config changes: give it a Code Interpreter so it does arithmetic in code, and tell it to check the raw text and the confidence score before it states a number. Here's the new prompt:

You answer finance questions about supplier invoices in the approval queue.
Use list_queue and get_invoice to look things up.

Before you state any invoice amount:
- Compare extracted.total (what the scanner read) with raw_total_text (what
  is printed on the invoice).
- If they don't match, or confidence is below 0.9, don't state the scanned
  amount as fact. Say what the invoice shows and mark it NEEDS HUMAN REVIEW.
- When you add up several invoices, use the code interpreter for the
  arithmetic and list the invoice IDs you included.

Keep answers short. Finance want the number, the invoice IDs, and anything
they need to check.
  1. Open Harness, select intake_agent, and choose Edit.

  2. Replace the system prompt with the one above.

  3. Under Tools, switch on Code interpreter tool. A picker appears underneath: choose AgentCore Code Interpreter, the built-in one AWS provides. Leave the Gateway tool as it is.

    Screenshot: Code interpreter tool switched on

    Tools section with Gateway and Code interpreter tool switched on, AgentCore Code Interpreter selected

  4. Save your changes.

cd $WORKSHOP/IntakeQA &&
agentcore add tool --harness intake_agent --type agentcore_code_interpreter --name calculator &&
$WORKSHOP/scripts/set-prompt.sh intake_agent $WORKSHOP/intake/system-prompt-v2.md &&
cat app/intake_agent/harness.json &&
agentcore deploy -y

set-prompt.sh writes the new prompt into systemPrompt in app/intake_agent/harness.json, because multi-line text in JSON is fiddly to edit by hand.

Saving creates version 2 of the harness. If it turned out to be worse, you could point an endpoint back at version 1.

Score live traffic from now on

To check the fix, set up online evaluation. It scores a sample of real sessions as they happen, so you don't have to run evaluations by hand. For the workshop, you'll score all of them.

  1. In Evaluations, open the Evaluation configurations tab and choose Create evaluation configuration.

  2. Evaluation name: intake_live. Leave Enable this evaluation configuration once created ticked.

    Screenshot: evaluation configuration details

    Create evaluation configuration: name intake_live, enable once created ticked

  3. Evaluators: tick Correctness and, under Custom evaluators, AmountCheck.

  4. Input data source: Define with an agent endpoint. Choose intake_agent and its DEFAULT endpoint.

  5. Sampling rate: change it from 10 to 100.

    Screenshot: data source and sampling

    Input data source with agent endpoint, Sampling rate 100 percent

  6. Leave Create default role under Permissions and choose Create evaluation configuration.

  7. Back in the Harness playground, ask the same four questions again, each in a new session.

Use the Console steps for this one. Like run eval, agentcore add online-eval expects a Runtime defined in the project, and a harness isn't one, so the CLI can't set this up for a harness yet. The console form creates the configuration and its role for you; choose IntakeQA_intake_agent as the agent.

Then ask the questions again from the terminal:

$WORKSHOP/scripts/intake-ask.sh intake_agent $WORKSHOP/intake/questions.txt

Scores arrive five to ten minutes after each session, which is about when wrap-up starts. Don't wait for them here; the wrap-up page shows you where to find them. The Lucky Star Print answers should now flag the invoice for review and pass AmountCheck.

In production you'd sample far less, typically 1 to 5 percent of sessions, because every scored session is another model call.

Optional: ask AgentCore to suggest a better prompt

AgentCore can read your evaluation results and propose an improved system prompt. This one is a terminal command on either path:

cd $WORKSHOP &&
agentcore run recommendation \
  -t system-prompt \
  -r intake_agent \
  -e AmountCheck \
  --prompt-file $WORKSHOP/intake/system-prompt-v2.md

Treat the output as a suggestion to review. In production you'd A/B test it against the current prompt before switching.

Kiro corner (optional)

Ask Kiro: "Read intake/questions.txt and intake/tools.json. Write five more finance questions, in the same format, that would catch the scanner misreading an amount." Then try them in the Playground, or add them to the file and run intake-ask.sh again.

Done when

  • You ran the built-in evaluators and can explain why they passed the wrong number.
  • AmountCheck fails the first Lucky Star Print answer.
  • After the version 2 prompt and the Code Interpreter, new answers flag the invoice for review.
  • Online evaluation is switched on. You'll check its scores at the start of wrap-up.

Stuck?

  • Jump to the finished state. In the code editor terminal, run the command below. It works on either path, asks before replacing anything you built in the console, and deploys version 2 of the prompt and the evaluator. Set up online evaluation in the console afterwards.

    $WORKSHOP/scripts/catch-up.sh intake
    
  • cd: no such file or directory: .../IntakeQA (terminal). The project doesn't exist yet: either the first step's Terminal box didn't finish, or you built the agent in the console. If you built it in the console, use the Console tabs for the rest of the scenario. Otherwise run the Terminal box in "Deploy the finance agent" again, or jump ahead with $WORKSHOP/scripts/catch-up.sh intake.

  • spanIds that do not exist in the provided data (terminal). A trace was still arriving when it was scored. This shows up most with Builtin.ToolSelectionAccuracy, which scores single tool calls. Wait a few minutes and run eval-harness.sh again for just that evaluator.

  • No sessions found when scoring. The traces aren't ready yet. Wait two minutes and try again, or widen the time range.

  • The evaluator won't save. The instructions must contain at least one placeholder for the level, such as {context}. Check that you pasted the whole text.

  • Online scores never appear. Check that the configuration is enabled. Results can take ten minutes to show up.