Skip to content

Failures

Every adapter's errors are normalized into one taxonomy, so calling code can decide what to do without knowing which engine produced the problem.

if not run.ok and run.failure is not None:
    run.failure.kind  # FailureKind.TIMEOUT — portable
    run.failure.retryable  # the condition may pass on its own
    run.safe_to_retry  # …and running this proposal again is harmless
    run.failure.native_code  # what the engine itself said

Gantry never retries

That is a deliberate boundary, not a missing feature. A retried write is a second attempt at something policy admitted once, and the layer whose job is to be the durable record of what ran is the wrong place to quietly run things twice. The flags below are advice for your code; nothing in Gantry acts on them.

retryable says one narrow thing

The condition that caused this failure may pass on its own. That is all. It is derived from the kind — FailureKind.transient is the single rule every adapter follows, and a test enforces it across the library so a timeout means the same thing on PostgreSQL as on BigQuery.

Transient Not transient
TIMEOUT, RESOURCE_ERROR, RESOURCE_EXHAUSTED everything else

Three kinds are left out on purpose. CONNECTOR_ERROR is a broker that is down or a sink table that does not exist, and the message rarely says which. ENGINE_ERROR and UNKNOWN mean Gantry could not tell what went wrong — and advice to retry on that basis has nothing behind it.

safe_to_retry answers the half that can destroy something

A transient condition does not make the work repeatable. Whether you may run the same proposal again depends on what it was and how far it got:

Operation Safe to run again
query Always, on a transient failure. Reads are idempotent.
materialize Only if nothing reached the engine.
batch, stream Only if nothing was submitted.

Materialization is create-only by design, and that is exactly what makes the retry unsafe: a run that reached the engine may have left its destination behind even though it failed, so running it again does not succeed — it comes back DESTINATION_EXISTS. A caller who saw retryable=True and simply tried again would turn a transient failure into one that looks permanent.

Submitted jobs are worse. Re-submitting a Flink job does not replace the first one, it adds a second.

run = await materialize(sql)

if run.safe_to_retry:
    run = await materialize(sql)  # nothing was created; this is a fresh attempt
elif run.failure and run.failure.retryable:
    ...  # transient, but the destination may exist —
    # check it, or write to a new destination

safe_to_retry returning False is not a prediction that the retry will fail. It is Gantry declining to promise the retry is harmless, which for a write is the answer that matters.

Transport-level retries

A dropped connection during describe() is unambiguously safe to retry and invisible to policy, so adapter-level retries for pure metadata reads would be reasonable. None are implemented today, and nothing that touches a governed operation will be retried this way — the distinction is whether policy admitted the call, not whether it looks cheap.

gantry.Failure dataclass

One normalized failure, keeping the engine's own words alongside.

kind and retryable are for code to act on; message is for a human. native_code, native_message and native preserve what the engine actually said, so normalizing never loses the detail needed to debug.

retryable says one narrow thing: the condition that caused this failure may pass on its own. It should equal kind.transient unless an adapter genuinely knows better about its own engine, and a test enforces that across the library so the flag means the same thing on every backend.

It does not say the work is safe to run again. A materialization that timed out halfway may have left its destination behind, and create-only means the retry fails with DESTINATION_EXISTS rather than succeeding. That question belongs to the run, which knows what kind of operation it was and how far it got: see Run.safe_to_retry.

Gantry itself never retries. A retried write is a second attempt at something policy admitted once, and deciding that belongs to the application, not to the layer whose job is to be the record of what ran.

Source code in gantry/failure.py
 83
 84
 85
 86
 87
 88
 89
 90
 91
 92
 93
 94
 95
 96
 97
 98
 99
100
101
102
103
104
105
106
107
108
109
110
111
112
@dataclass(frozen=True, slots=True)
class Failure:
    """One normalized failure, keeping the engine's own words alongside.

    `kind` and `retryable` are for code to act on; `message` is for a human.
    `native_code`, `native_message` and `native` preserve what the engine
    actually said, so normalizing never loses the detail needed to debug.

    `retryable` says one narrow thing: **the condition that caused this failure
    may pass on its own.** It should equal `kind.transient` unless an adapter
    genuinely knows better about its own engine, and a test enforces that across
    the library so the flag means the same thing on every backend.

    It does not say the work is safe to run again. A materialization that timed
    out halfway may have left its destination behind, and create-only means the
    retry fails with `DESTINATION_EXISTS` rather than succeeding. That question
    belongs to the run, which knows what kind of operation it was and how far it
    got: see `Run.safe_to_retry`.

    Gantry itself never retries. A retried write is a second attempt at
    something policy admitted once, and deciding that belongs to the
    application, not to the layer whose job is to be the record of what ran.
    """

    kind: FailureKind
    retryable: bool
    message: str
    native_code: str | None = None
    native_message: str | None = None
    native: Mapping[str, object] = field(default_factory=dict)

gantry.FailureKind

Bases: StrEnum

A portable reason an execution did not produce a trustworthy result.

Normalized across engines so a caller can branch on the kind rather than parse an engine's message. The distinctions that matter most are refusals before execution (POLICY_REJECTED, UNSUPPORTED_POLICY_REQUIREMENT), failures during it (ENGINE_ERROR, TIMEOUT), and a run that finished but cannot be believed (VERIFICATION_FAILED).

Source code in gantry/failure.py
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
class FailureKind(StrEnum):
    """A portable reason an execution did not produce a trustworthy result.

    Normalized across engines so a caller can branch on the kind rather than
    parse an engine's message. The distinctions that matter most are refusals
    before execution (`POLICY_REJECTED`,
    `UNSUPPORTED_POLICY_REQUIREMENT`), failures during it (`ENGINE_ERROR`,
    `TIMEOUT`), and a run that finished but cannot be believed
    (`VERIFICATION_FAILED`).
    """

    VALIDATION_ERROR = "VALIDATION_ERROR"
    POLICY_REJECTED = "POLICY_REJECTED"
    UNSUPPORTED_POLICY_REQUIREMENT = "UNSUPPORTED_POLICY_REQUIREMENT"
    INPUT_NOT_ALLOWED = "INPUT_NOT_ALLOWED"
    OUTPUT_NOT_ALLOWED = "OUTPUT_NOT_ALLOWED"
    SOURCE_NOT_ALLOWED = "SOURCE_NOT_ALLOWED"
    DESTINATION_NOT_ALLOWED = "DESTINATION_NOT_ALLOWED"
    DESTINATION_EXISTS = "DESTINATION_EXISTS"
    OPERATION_NOT_ALLOWED = "OPERATION_NOT_ALLOWED"
    SUBMISSION_ERROR = "SUBMISSION_ERROR"
    AUTH_ERROR = "AUTH_ERROR"
    OBJECT_NOT_FOUND = "OBJECT_NOT_FOUND"
    SYNTAX_ERROR = "SYNTAX_ERROR"
    CONNECTOR_ERROR = "CONNECTOR_ERROR"
    RESOURCE_ERROR = "RESOURCE_ERROR"
    RESOURCE_EXHAUSTED = "RESOURCE_EXHAUSTED"
    COST_LIMIT_EXCEEDED = "COST_LIMIT_EXCEEDED"
    TIMEOUT = "TIMEOUT"
    ENGINE_ERROR = "ENGINE_ERROR"
    USER_CODE_ERROR = "USER_CODE_ERROR"
    DATA_ERROR = "DATA_ERROR"
    CANCELLED = "CANCELLED"
    VERIFICATION_FAILED = "VERIFICATION_FAILED"
    UNSUPPORTED_VERIFICATION = "UNSUPPORTED_VERIFICATION"
    VERIFICATION_UNSUPPORTED = "VERIFICATION_UNSUPPORTED"
    VERIFICATION_CONFLICT = "VERIFICATION_CONFLICT"
    UNKNOWN = "UNKNOWN"

    @property
    def transient(self) -> bool:
        """Whether this kind describes a condition that may pass on its own.

        The rule every adapter should agree with, written once. Before this
        existed each adapter carried its own copy — BigQuery tested
        `kind in {TIMEOUT, RESOURCE_EXHAUSTED}`, Snowflake `kind is TIMEOUT`,
        Flink an if-chain — and nothing made a new adapter agree with the old
        ones.

        The set is deliberately small and conservative, because the cost of the
        two mistakes is not symmetric. Calling a permanent failure transient
        invites a caller to retry something that will never succeed; calling a
        transient one permanent only costs them an attempt they could have made.

        Three kinds are left out on purpose. `CONNECTOR_ERROR` is a broker that
        is down *or* a sink table that does not exist, and the message rarely
        says which. `ENGINE_ERROR` and `UNKNOWN` mean Gantry could not tell what
        went wrong, and advising a retry on that basis is advice with nothing
        behind it.
        """
        return self in _TRANSIENT

transient property

transient: bool

Whether this kind describes a condition that may pass on its own.

The rule every adapter should agree with, written once. Before this existed each adapter carried its own copy — BigQuery tested kind in {TIMEOUT, RESOURCE_EXHAUSTED}, Snowflake kind is TIMEOUT, Flink an if-chain — and nothing made a new adapter agree with the old ones.

The set is deliberately small and conservative, because the cost of the two mistakes is not symmetric. Calling a permanent failure transient invites a caller to retry something that will never succeed; calling a transient one permanent only costs them an attempt they could have made.

Three kinds are left out on purpose. CONNECTOR_ERROR is a broker that is down or a sink table that does not exist, and the message rarely says which. ENGINE_ERROR and UNKNOWN mean Gantry could not tell what went wrong, and advising a retry on that basis is advice with nothing behind it.

gantry.SubmissionError

Bases: Exception

Structured rejection or submission failure from :meth:ControlPlane.submit.

Source code in gantry/runtime.py
42
43
44
45
46
47
48
class SubmissionError(Exception):
    """Structured rejection or submission failure from :meth:`ControlPlane.submit`."""

    def __init__(self, result: Result) -> None:
        self.result = result
        message = result.failure.message if result.failure is not None else result.status.value
        super().__init__(message)