Failures¶
Every adapter's errors are normalized into one taxonomy, so calling code can decide what to do without knowing which engine produced the problem.
if not run.ok and run.failure is not None:
run.failure.kind # FailureKind.TIMEOUT — portable
run.failure.retryable # the condition may pass on its own
run.safe_to_retry # …and running this proposal again is harmless
run.failure.native_code # what the engine itself said
Gantry never retries¶
That is a deliberate boundary, not a missing feature. A retried write is a second attempt at something policy admitted once, and the layer whose job is to be the durable record of what ran is the wrong place to quietly run things twice. The flags below are advice for your code; nothing in Gantry acts on them.
retryable says one narrow thing¶
The condition that caused this failure may pass on its own. That is all. It
is derived from the kind — FailureKind.transient is the single rule every
adapter follows, and a test enforces it across the library so a timeout means
the same thing on PostgreSQL as on BigQuery.
| Transient | Not transient |
|---|---|
TIMEOUT, RESOURCE_ERROR, RESOURCE_EXHAUSTED |
everything else |
Three kinds are left out on purpose. CONNECTOR_ERROR is a broker that is down
or a sink table that does not exist, and the message rarely says which.
ENGINE_ERROR and UNKNOWN mean Gantry could not tell what went wrong — and
advice to retry on that basis has nothing behind it.
safe_to_retry answers the half that can destroy something¶
A transient condition does not make the work repeatable. Whether you may run the same proposal again depends on what it was and how far it got:
| Operation | Safe to run again |
|---|---|
| query | Always, on a transient failure. Reads are idempotent. |
| materialize | Only if nothing reached the engine. |
| batch, stream | Only if nothing was submitted. |
Materialization is create-only by design, and that is exactly what makes the
retry unsafe: a run that reached the engine may have left its destination behind
even though it failed, so running it again does not succeed — it comes back
DESTINATION_EXISTS. A caller who saw retryable=True and simply tried again
would turn a transient failure into one that looks permanent.
Submitted jobs are worse. Re-submitting a Flink job does not replace the first one, it adds a second.
run = await materialize(sql)
if run.safe_to_retry:
run = await materialize(sql) # nothing was created; this is a fresh attempt
elif run.failure and run.failure.retryable:
... # transient, but the destination may exist —
# check it, or write to a new destination
safe_to_retry returning False is not a prediction that the retry will fail.
It is Gantry declining to promise the retry is harmless, which for a write is
the answer that matters.
Transport-level retries¶
A dropped connection during describe() is unambiguously safe to retry and
invisible to policy, so adapter-level retries for pure metadata reads would be
reasonable. None are implemented today, and nothing that touches a governed
operation will be retried this way — the distinction is whether policy admitted
the call, not whether it looks cheap.
gantry.Failure
dataclass
¶
One normalized failure, keeping the engine's own words alongside.
kind and retryable are for code to act on; message is for a human.
native_code, native_message and native preserve what the engine
actually said, so normalizing never loses the detail needed to debug.
retryable says one narrow thing: the condition that caused this failure
may pass on its own. It should equal kind.transient unless an adapter
genuinely knows better about its own engine, and a test enforces that across
the library so the flag means the same thing on every backend.
It does not say the work is safe to run again. A materialization that timed
out halfway may have left its destination behind, and create-only means the
retry fails with DESTINATION_EXISTS rather than succeeding. That question
belongs to the run, which knows what kind of operation it was and how far it
got: see Run.safe_to_retry.
Gantry itself never retries. A retried write is a second attempt at something policy admitted once, and deciding that belongs to the application, not to the layer whose job is to be the record of what ran.
Source code in gantry/failure.py
83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 | |
gantry.FailureKind
¶
Bases: StrEnum
A portable reason an execution did not produce a trustworthy result.
Normalized across engines so a caller can branch on the kind rather than
parse an engine's message. The distinctions that matter most are refusals
before execution (POLICY_REJECTED,
UNSUPPORTED_POLICY_REQUIREMENT), failures during it (ENGINE_ERROR,
TIMEOUT), and a run that finished but cannot be believed
(VERIFICATION_FAILED).
Source code in gantry/failure.py
11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 | |
transient
property
¶
transient: bool
Whether this kind describes a condition that may pass on its own.
The rule every adapter should agree with, written once. Before this
existed each adapter carried its own copy — BigQuery tested
kind in {TIMEOUT, RESOURCE_EXHAUSTED}, Snowflake kind is TIMEOUT,
Flink an if-chain — and nothing made a new adapter agree with the old
ones.
The set is deliberately small and conservative, because the cost of the two mistakes is not symmetric. Calling a permanent failure transient invites a caller to retry something that will never succeed; calling a transient one permanent only costs them an attempt they could have made.
Three kinds are left out on purpose. CONNECTOR_ERROR is a broker that
is down or a sink table that does not exist, and the message rarely
says which. ENGINE_ERROR and UNKNOWN mean Gantry could not tell what
went wrong, and advising a retry on that basis is advice with nothing
behind it.
gantry.SubmissionError
¶
Bases: Exception
Structured rejection or submission failure from :meth:ControlPlane.submit.
Source code in gantry/runtime.py
42 43 44 45 46 47 48 | |