Skip to content

Evidence and runs

A result says a run was accepted. Evidence says why — in enough detail that someone who was not there, and who does not have the agent conversation that produced the work, can decide whether they agree.

Observations

Three sources, kept apart because they are trusted differently. The engine's account of its own execution is not the same kind of fact as a measurement Gantry took at the destination.

gantry.Observation dataclass

One bounded measurement, with the source that produced it.

A scalar, deliberately. "The destination has 1,190,432 rows" is an observation; the rows themselves are not.

Source code in gantry/evidence.py
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
@dataclass(frozen=True, slots=True)
class Observation:
    """One bounded measurement, with the source that produced it.

    A scalar, deliberately. "The destination has 1,190,432 rows" is an
    observation; the rows themselves are not.
    """

    name: str
    value: object
    source: ObservationSource
    unit: str | None = None
    observed_at: datetime = field(default_factory=lambda: datetime.now(UTC))

    def __post_init__(self) -> None:
        if not self.name.strip():
            raise ValueError("observation name must not be empty")

    def as_dict(self) -> dict[str, object]:
        payload: dict[str, object] = {
            "name": self.name,
            "value": _plain(self.value),
            "source": self.source.value,
            "observed_at": self.observed_at.isoformat(),
        }
        if self.unit is not None:
            payload["unit"] = self.unit
        return payload

gantry.ObservationSource

Bases: StrEnum

Where an observation came from, and therefore how far to trust it.

ENGINE is the engine's account of its own work. OUTPUT is a measurement of what was produced, taken by Gantry rather than reported by the thing that produced it. GANTRY is what the control plane itself did — the checks requested, the decision reached.

Source code in gantry/evidence.py
31
32
33
34
35
36
37
38
39
40
41
42
class ObservationSource(StrEnum):
    """Where an observation came from, and therefore how far to trust it.

    `ENGINE` is the engine's account of its own work. `OUTPUT` is a measurement
    of what was produced, taken by Gantry rather than reported by the thing that
    produced it. `GANTRY` is what the control plane itself did — the checks
    requested, the decision reached.
    """

    ENGINE = "engine"
    OUTPUT = "output"
    GANTRY = "gantry"

gantry.CheckSource

Bases: StrEnum

Who committed Gantry to a verification requirement.

Source code in gantry/verifier.py
16
17
18
19
20
class CheckSource(StrEnum):
    """Who committed Gantry to a verification requirement."""

    TRUSTED = "trusted"
    AGENT = "agent"

gantry.VerificationSource module-attribute

VerificationSource = CheckSource

The bundle

gantry.EvidenceBundle dataclass

The record of one governed run: what ran, what was seen, what was decided.

Serializable on purpose. The point of evidence is that it outlives the process that gathered it, so everything here survives json.dumps.

Source code in gantry/evidence.py
 75
 76
 77
 78
 79
 80
 81
 82
 83
 84
 85
 86
 87
 88
 89
 90
 91
 92
 93
 94
 95
 96
 97
 98
 99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
@dataclass(frozen=True, slots=True)
class EvidenceBundle:
    """The record of one governed run: what ran, what was seen, what was decided.

    Serializable on purpose. The point of evidence is that it outlives the
    process that gathered it, so everything here survives `json.dumps`.
    """

    run_id: str
    engine: str
    operation: str
    decision: str
    native_execution_id: str | None = None
    proposal_hash: str | None = None
    inputs: tuple[str, ...] = ()
    outputs: tuple[str, ...] = ()
    started_at: datetime | None = None
    finished_at: datetime | None = None
    execution: Mapping[str, object] = field(default_factory=dict)
    proposal: Mapping[str, object] = field(default_factory=dict)
    observations: tuple[Observation, ...] = ()
    checks: tuple[CheckResult, ...] = ()

    @property
    def trusted_checks(self) -> tuple[CheckResult, ...]:
        return tuple(check for check in self.checks if check.source == CheckSource.TRUSTED)

    @property
    def agent_checks(self) -> tuple[CheckResult, ...]:
        return tuple(check for check in self.checks if check.source == CheckSource.AGENT)

    @property
    def verification(self) -> VerificationResult:
        return VerificationResult(ok=all(check.ok for check in self.checks), checks=self.checks)

    @property
    def duration_ms(self) -> int | None:
        if self.started_at is None or self.finished_at is None:
            return None
        return int((self.finished_at - self.started_at).total_seconds() * 1000)

    def observation(self, name: str) -> Observation | None:
        """The most recent observation with this name, if one was made."""
        return next(
            (item for item in reversed(self.observations) if item.name == name),
            None,
        )

    def as_dict(self) -> dict[str, object]:
        return {
            "run_id": self.run_id,
            "engine": self.engine,
            "operation": self.operation,
            "decision": self.decision,
            "native_execution_id": self.native_execution_id,
            "proposal_hash": self.proposal_hash,
            "inputs": list(self.inputs),
            "outputs": list(self.outputs),
            "started_at": None if self.started_at is None else self.started_at.isoformat(),
            "finished_at": None if self.finished_at is None else self.finished_at.isoformat(),
            "duration_ms": self.duration_ms,
            "execution": {key: _plain(value) for key, value in self.execution.items()},
            "proposal": {key: _plain(value) for key, value in self.proposal.items()},
            "observations": [item.as_dict() for item in self.observations],
            "checks": [_check_as_dict(check) for check in self.checks],
            "trusted_checks": [_check_as_dict(check) for check in self.trusted_checks],
            "agent_checks": [_check_as_dict(check) for check in self.agent_checks],
            "verification": {
                "passed": all(check.ok for check in self.checks),
                "checks": [_check_as_dict(check) for check in self.checks],
            },
        }

    def to_json(self, *, indent: int | None = None) -> str:
        return json.dumps(self.as_dict(), indent=indent, sort_keys=False)

observation

observation(name: str) -> Observation | None

The most recent observation with this name, if one was made.

Source code in gantry/evidence.py
116
117
118
119
120
121
def observation(self, name: str) -> Observation | None:
    """The most recent observation with this name, if one was made."""
    return next(
        (item for item in reversed(self.observations) if item.name == name),
        None,
    )

References, not data

An observation is a name, a scalar and a source, so a bundle stays small however large the output it describes. The row count of a billion-row table is one integer. Gantry is not the data plane, and evidence is not a copy of the data.

The same checks on both paths

gantry.verify serves a query and a materialization. A materialization is checked against the destination it created; a query against the rows it returned, described as a table so one check means one thing.

checks = [gantry.verify.row_count(min=1), gantry.verify.required_columns(["id"])]

await db.query(schemas=["analytics"], checks=checks)(sql)
await db.materialize(sources=["analytics.*"], destinations=["reporting.*"], checks=checks)(sql)

Configured checks= are trusted application commitments. A caller may add task-specific checks at invocation time; the merge is additive and Gantry assigns provenance itself:

query = db.query(
    schemas=["analytics"],
    checks=[gantry.verify.row_count(max=1000)],
)

result = await query(
    sql,
    verify=[
        gantry.verify.not_empty(),
        gantry.verify.required_columns(["customer_id", "revenue"]),
    ],
)

query.tool() exposes verify as an allowlisted declarative JSON schema. An unsupported check, attempted source override, or executable callback is rejected. Statically contradictory count bounds return VERIFICATION_CONFLICT before the engine is started.

Two cases cannot be the same, and both fail closed rather than pretending:

  • destination_exists and output_exists ask about something a query never creates, so on a query they report themselves unsupported. Answering them against the result set would make them trivially true, and a caller would believe a destination had been checked.
  • A result truncated by max_rows describes the rows returned, not the rows matched, so the count-based checks are unsupported there. Counting what came back would measure the policy rather than the data.

Both arrive as FailureKind.UNSUPPORTED_VERIFICATION, which is distinct from a check that ran and failed.

Durable runs

Evidence is attached to a run, which is the durable record of the whole operation. See Runs.