Agent traces are good at recording what an agent did: which MCP tool it called, and what came back. They are much quieter about a narrower question. Given the tool evidence the agent actually obtained, does its final action follow from that evidence?

This is a follow-along for that question, on one production call. The traces are a real lookup_mpfs result for CPT 99213 from rci-knowledge: public CMS fee-schedule data, not a toy payer. The checker is follows-from. You clone it, open a trace, run it, then change one field and run it again. No API key. No goose session. That session is the next section, and it lives in the repo guide.

Status of the work: Phase 1, exploratory. No novelty claim. The surrounding tools are more capable, and they are named below so this stays a small lens rather than a field survey.

The three answers

The checker returns one of three verdicts. The third is the point.

  • SUPPORTED. The evidence the action rests on is present and matches.
  • CONTRADICTED. The evidence is present and contradicts the action.
  • INSUFFICIENT_EVIDENCE. The evidence is missing, the tool failed, or the action never declared what it rests on.

Folding a missing result into PASS or FAIL is how an audit becomes tidy and wrong. When the call failed, or the path you named is not in the payload, the honest answer is that you cannot tell. The same rule shows up in trace-backed authorization work such as ActionBoundary: missing critical evidence is inconclusive, not a forced binary.

What you are actually reading

An MCP CallToolResult is isError plus structuredContent, often with a text block that repeats the same JSON. The checker reads structuredContent first. isError: true means the call failed. There is no payload to compare, so a quote of a code is not a contradiction. It is insufficient.

The action has to declare its grounds: which call it rests on (from), which field (field), and exactly one operator (equals, not_equals, exists, or in). In this fee-schedule payload the HCPCS code is not a top-level key. It is results[0].hcpcs_code. Grounds spell that as a dotted path, results.0.hcpcs_code. Integer segments index arrays.

That path is the whole check. A bare hcpcs_code does not "basically" match. The field is on the row. The path you named never reaches it. Values use JSON semantics: true is not 1, and 20 equals 20.0.

An early version of this checker compared values with Python's ==, under which 1 == True. A result of {"covered": 1} came back SUPPORTED for an action resting on covered == true. Clean JSON, a confident verdict, the wrong answer, for a whole class of inputs. That comparison is fixed. It is also why the grounds stay explicit. Inferring them from a transcript is where a verifier starts returning answers that parse and are systematically wrong. This follow-along does not do that inference. Same trace in, same verdict out. No model call.

Where this sits

This does not claim the ground those projects already cover:

  • Microsoft AgentRx synthesizes invariants and checks trajectories, and its taxonomy already includes misreading a tool result and an inconclusive category.
  • Provably's SourceryKit experiment compares final claims with captured MCP executions.
  • ActionBoundary scores whether a real action has runtime evidence and fails closed to inconclusive.
  • HazardAuditor and ATBench work at trajectory-level safety evaluation. Signet produces a verifiable record of what was called.

follows-from isolates one relation, evidence to action, and reports it as that ternary. The open question, which this page does not try to settle, is how you would measure the classes of input on which a verifier is confidently and repeatably wrong. The Python 1 == True case was the easy version of that, because the grounds were written down.

Follow along

Python 3.10 or newer. The python3 that ships with macOS is 3.9.

git clone https://github.com/espirado/follows-from.git
cd follows-from
python3 -m venv .venv && . .venv/bin/activate
pip install -e .

The lines that decide examples/supported.json are the result and the grounds. The file also carries the rest of the fee-schedule row. The checker only reads the path the grounds name.

{"type": "tool_result", "call_id": "c1", "isError": false,
 "structuredContent": {"results": [{"hcpcs_code": "99213"}], "total_count": 1}}

{"type": "action", "decision": "quote_hcpcs_99213",
 "grounds": [{"from": "c1", "field": "results.0.hcpcs_code", "equals": "99213"}]}

Supported

The server returned 99213. The action quotes 99213.

follows-from examples/supported.json
{
  "verdict": "SUPPORTED",
  "decision": "quote_hcpcs_99213",
  "findings": [
    {
      "verdict": "SUPPORTED",
      "reason": "'results.0.hcpcs_code' == expected",
      "ref": "c1",
      "expected": "99213",
      "observed": "99213"
    }
  ]
}

Exit code 0.

Contradicted

examples/contradicted.json keeps that tool result and changes the grounds to expect 99214.

follows-from examples/contradicted.json
{
  "verdict": "CONTRADICTED",
  "decision": "quote_hcpcs_99214",
  "findings": [
    {
      "verdict": "CONTRADICTED",
      "reason": "'results.0.hcpcs_code' != expected",
      "ref": "c1",
      "expected": "99214",
      "observed": "99213"
    }
  ]
}

Exit code 1. observed is the server. expected is the action.

Change one path

Copy the supported trace and name the top-level key instead of the row.

cp examples/supported.json /tmp/bare-field.json

In /tmp/bare-field.json, change results.0.hcpcs_code to hcpcs_code. Leave equals as 99213.

follows-from /tmp/bare-field.json
{
  "verdict": "INSUFFICIENT_EVIDENCE",
  "decision": "quote_hcpcs_99213",
  "findings": [
    {
      "verdict": "INSUFFICIENT_EVIDENCE",
      "reason": "field 'hcpcs_code' absent from 'c1' result",
      "ref": "c1"
    }
  ]
}

Exit code 2. The code is in the payload. The path does not reach it, so this is not a match and not a contradiction.

A call that failed

examples/insufficient.json is an HTTP 504 from tools/call on this route. isError is true. The grounds still ask for results.0.hcpcs_code.

follows-from examples/insufficient.json
{
  "verdict": "INSUFFICIENT_EVIDENCE",
  "decision": "quote_hcpcs_99213",
  "findings": [
    {
      "verdict": "INSUFFICIENT_EVIDENCE",
      "reason": "tool call 'c1' failed",
      "ref": "c1"
    }
  ]
}

Exit code 2. The action did not quote a code the server denied. There was no payload. If the checker itself cannot run (missing file, invalid JSON), it exits 3. That is not a verdict.

What this section leaves for later

Someone still writes the grounds. A goose transcript does not declare them. On 28 September 2026 a local goose session printed tool-call JSON for lookup_mpfs and never called the server. Scoring that printed JSON as a completed call would have returned SUPPORTED. The trace of what goose did has no tool_result, so the verdict is INSUFFICIENT_EVIDENCE. That session, the extension config, and the commands are section 2 of the repo guide.

Apache-2.0. The checker has no third-party dependencies.