Evidence and runs¶
A result says a run was accepted. Evidence says why — in enough detail that someone who was not there, and who does not have the agent conversation that produced the work, can decide whether they agree.
Observations¶
Three sources, kept apart because they are trusted differently. The engine's account of its own execution is not the same kind of fact as a measurement Gantry took at the destination.
gantry.Observation
dataclass
¶
One bounded measurement, with the source that produced it.
A scalar, deliberately. "The destination has 1,190,432 rows" is an observation; the rows themselves are not.
Source code in gantry/evidence.py
45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 | |
gantry.ObservationSource
¶
Bases: StrEnum
Where an observation came from, and therefore how far to trust it.
ENGINE is the engine's account of its own work. OUTPUT is a measurement
of what was produced, taken by Gantry rather than reported by the thing that
produced it. GANTRY is what the control plane itself did — the checks
requested, the decision reached.
Source code in gantry/evidence.py
31 32 33 34 35 36 37 38 39 40 41 42 | |
gantry.CheckSource
¶
Bases: StrEnum
Who committed Gantry to a verification requirement.
Source code in gantry/verifier.py
16 17 18 19 20 | |
The bundle¶
gantry.EvidenceBundle
dataclass
¶
The record of one governed run: what ran, what was seen, what was decided.
Serializable on purpose. The point of evidence is that it outlives the
process that gathered it, so everything here survives json.dumps.
Source code in gantry/evidence.py
75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 | |
observation
¶
observation(name: str) -> Observation | None
The most recent observation with this name, if one was made.
Source code in gantry/evidence.py
116 117 118 119 120 121 | |
References, not data
An observation is a name, a scalar and a source, so a bundle stays small however large the output it describes. The row count of a billion-row table is one integer. Gantry is not the data plane, and evidence is not a copy of the data.
The same checks on both paths¶
gantry.verify serves a query and a materialization. A materialization is
checked against the destination it created; a query against the rows it
returned, described as a table so one check means one thing.
checks = [gantry.verify.row_count(min=1), gantry.verify.required_columns(["id"])]
await db.query(schemas=["analytics"], checks=checks)(sql)
await db.materialize(sources=["analytics.*"], destinations=["reporting.*"], checks=checks)(sql)
Configured checks= are trusted application commitments. A caller may add
task-specific checks at invocation time; the merge is additive and Gantry
assigns provenance itself:
query = db.query(
schemas=["analytics"],
checks=[gantry.verify.row_count(max=1000)],
)
result = await query(
sql,
verify=[
gantry.verify.not_empty(),
gantry.verify.required_columns(["customer_id", "revenue"]),
],
)
query.tool() exposes verify as an allowlisted declarative JSON schema. An
unsupported check, attempted source override, or
executable callback is rejected. Statically contradictory count bounds return
VERIFICATION_CONFLICT before the engine is started.
Two cases cannot be the same, and both fail closed rather than pretending:
destination_existsandoutput_existsask about something a query never creates, so on a query they report themselves unsupported. Answering them against the result set would make them trivially true, and a caller would believe a destination had been checked.- A result truncated by
max_rowsdescribes the rows returned, not the rows matched, so the count-based checks are unsupported there. Counting what came back would measure the policy rather than the data.
Both arrive as FailureKind.UNSUPPORTED_VERIFICATION, which is distinct from a
check that ran and failed.
Durable runs¶
Evidence is attached to a run, which is the durable record of the whole operation. See Runs.