Agent traces are good at recording what an agent did: which MCP tool it called, and what came back. They are much quieter about a narrower question. Given the tool evidence the agent actually obtained, does its final action follow from that evidence?
This is a follow-along for that question, on one production call. The traces are a real lookup_mpfs result for CPT 99213 from rci-knowledge: public CMS fee-schedule data, not a toy payer. The checker is follows-from. You clone it, open a trace, run it, then change one field and run it again. No API key. No goose session. That session is the next section, and it lives in the repo guide.
Status of the work: Phase 1, exploratory. No novelty claim. The surrounding tools are more capable, and they are named below so this stays a small lens rather than a field survey.
The checker returns one of three verdicts. The third is the point.
Folding a missing result into PASS or FAIL is how an audit becomes tidy and wrong. When the call failed, or the path you named is not in the payload, the honest answer is that you cannot tell. The same rule shows up in trace-backed authorization work such as ActionBoundary: missing critical evidence is inconclusive, not a forced binary.
An MCP CallToolResult is isError plus structuredContent, often with a text block that repeats the same JSON. The checker reads structuredContent first. isError: true means the call failed. There is no payload to compare, so a quote of a code is not a contradiction. It is insufficient.
The action has to declare its grounds: which call it rests on (from), which field (field), and exactly one operator (equals, not_equals, exists, or in). In this fee-schedule payload the HCPCS code is not a top-level key. It is results[0].hcpcs_code. Grounds spell that as a dotted path, results.0.hcpcs_code. Integer segments index arrays.
That path is the whole check. A bare hcpcs_code does not "basically" match. The field is on the row. The path you named never reaches it. Values use JSON semantics: true is not 1, and 20 equals 20.0.
An early version of this checker compared values with Python's ==, under which 1 == True. A result of {"covered": 1} came back SUPPORTED for an action resting on covered == true. Clean JSON, a confident verdict, the wrong answer, for a whole class of inputs. That comparison is fixed. It is also why the grounds stay explicit. Inferring them from a transcript is where a verifier starts returning answers that parse and are systematically wrong. This follow-along does not do that inference. Same trace in, same verdict out. No model call.
This does not claim the ground those projects already cover:
follows-from isolates one relation, evidence to action, and reports it as that ternary. The open question, which this page does not try to settle, is how you would measure the classes of input on which a verifier is confidently and repeatably wrong. The Python 1 == True case was the easy version of that, because the grounds were written down.
Python 3.10 or newer. The python3 that ships with macOS is 3.9.
git clone https://github.com/espirado/follows-from.git
cd follows-from
python3 -m venv .venv && . .venv/bin/activate
pip install -e .
The lines that decide examples/supported.json are the result and the grounds. The file also carries the rest of the fee-schedule row. The checker only reads the path the grounds name.
{"type": "tool_result", "call_id": "c1", "isError": false,
"structuredContent": {"results": [{"hcpcs_code": "99213"}], "total_count": 1}}
{"type": "action", "decision": "quote_hcpcs_99213",
"grounds": [{"from": "c1", "field": "results.0.hcpcs_code", "equals": "99213"}]}
The server returned 99213. The action quotes 99213.
follows-from examples/supported.json
{
"verdict": "SUPPORTED",
"decision": "quote_hcpcs_99213",
"findings": [
{
"verdict": "SUPPORTED",
"reason": "'results.0.hcpcs_code' == expected",
"ref": "c1",
"expected": "99213",
"observed": "99213"
}
]
}
Exit code 0.
examples/contradicted.json keeps that tool result and changes the grounds to expect 99214.
follows-from examples/contradicted.json
{
"verdict": "CONTRADICTED",
"decision": "quote_hcpcs_99214",
"findings": [
{
"verdict": "CONTRADICTED",
"reason": "'results.0.hcpcs_code' != expected",
"ref": "c1",
"expected": "99214",
"observed": "99213"
}
]
}
Exit code 1. observed is the server. expected is the action.
Copy the supported trace and name the top-level key instead of the row.
cp examples/supported.json /tmp/bare-field.json
In /tmp/bare-field.json, change results.0.hcpcs_code to hcpcs_code. Leave equals as 99213.
follows-from /tmp/bare-field.json
{
"verdict": "INSUFFICIENT_EVIDENCE",
"decision": "quote_hcpcs_99213",
"findings": [
{
"verdict": "INSUFFICIENT_EVIDENCE",
"reason": "field 'hcpcs_code' absent from 'c1' result",
"ref": "c1"
}
]
}
Exit code 2. The code is in the payload. The path does not reach it, so this is not a match and not a contradiction.
examples/insufficient.json is an HTTP 504 from tools/call on this route. isError is true. The grounds still ask for results.0.hcpcs_code.
follows-from examples/insufficient.json
{
"verdict": "INSUFFICIENT_EVIDENCE",
"decision": "quote_hcpcs_99213",
"findings": [
{
"verdict": "INSUFFICIENT_EVIDENCE",
"reason": "tool call 'c1' failed",
"ref": "c1"
}
]
}
Exit code 2. The action did not quote a code the server denied. There was no payload. If the checker itself cannot run (missing file, invalid JSON), it exits 3. That is not a verdict.
Someone still writes the grounds. A goose transcript does not declare them. On 28 September 2026 a local goose session printed tool-call JSON for lookup_mpfs and never called the server. Scoring that printed JSON as a completed call would have returned SUPPORTED. The trace of what goose did has no tool_result, so the verdict is INSUFFICIENT_EVIDENCE. That session, the extension config, and the commands are section 2 of the repo guide.
Apache-2.0. The checker has no third-party dependencies.