Per-call timeouts are the default everywhere and they compose badly: three hops with a two-second timeout each produce a six-second failure for a user who left after three. A deadline belongs to the request, travels with it, and is spent down — which turns a chain of independent guesses into one budget with admission control. This guide implements that, as the timing mechanism behind fallback routing for geospatial queries.
When to Use This Approach
Propagate a deadline whenever a request crosses more than one boundary, which for a spatial agent is essentially always: resolution, retrieval, geometry, and possibly a fallback ladder inside each.
| Pattern | Effect on a slow failure | Use |
|---|---|---|
| Per-call timeout | Costs the sum of every timeout | Only for a single-hop request |
| Shared deadline | Costs one budget, total | The default |
| Deadline plus admission | Costs less than one budget | Where cheap fallbacks exist |
| No timeout | Costs whatever the slowest dependency does | Never |
The third row is what makes a fallback ladder work. Without admission control the expensive first rung consumes the whole budget before failing, and the cheap cached rung that would have satisfied the user never runs.
Implementation
The deadline is a small object created once per request, passed down, and consulted before every call.
import logging
import time
from dataclasses import dataclass
from typing import Optional
log = logging.getLogger("deadline")
class DeadlineExceeded(TimeoutError):
"""The request budget is spent; nothing further may be started."""
@dataclass(frozen=True)
class Deadline:
started_monotonic: float
budget_s: float
label: str = "request"
@classmethod
def start(cls, budget_s: float, label: str = "request") -> "Deadline":
if budget_s <= 0:
raise ValueError("deadline budget must be positive")
return cls(time.monotonic(), float(budget_s), label)
def remaining_s(self) -> float:
return self.budget_s - (time.monotonic() - self.started_monotonic)
def expired(self) -> bool:
return self.remaining_s() <= 0
def check(self) -> None:
if self.expired():
raise DeadlineExceeded(f"{self.label} budget of {self.budget_s:.1f}s is spent")
def sub(self, fraction: float, label: str) -> "Deadline":
"""A child deadline for one stage — never longer than what remains."""
left = self.remaining_s()
if left <= 0:
raise DeadlineExceeded(f"{self.label} budget is spent before {label}")
return Deadline(time.monotonic(), min(left, max(0.0, self.budget_s * fraction)), label)
Using a monotonic clock rather than wall time is not a detail: a wall clock can move backwards under a time synchronisation and produce a deadline that never expires or expires immediately, and the failure is rare enough to be baffling when it happens.
Admission control is the second half, and the part most implementations omit.
def admit(deadline: Deadline, typical_s: float, name: str,
safety: float = 0.5) -> bool:
"""Should this call be started at all, given what is left?"""
left = deadline.remaining_s()
if left <= 0:
log.info("skipping %s: budget already spent", name)
return False
if left < typical_s * safety:
log.info("skipping %s: %.2fs left, typically needs %.2fs", name, left, typical_s)
return False
return True
def call_with_deadline(fn, deadline: Deadline, typical_s: float, name: str, *args):
"""Run a dependency with whatever time remains, or decline to start it."""
if not admit(deadline, typical_s, name):
raise DeadlineExceeded(f"no time left to attempt {name}")
timeout = max(0.05, deadline.remaining_s())
started = time.monotonic()
try:
return fn(*args, timeout=timeout)
finally:
log.debug("%s took %.3fs of %.3fs remaining", name,
time.monotonic() - started, timeout)
The safety factor is what makes admission useful rather than merely conservative. Requiring the full typical duration would decline calls that frequently finish faster; requiring half of it declines only the attempts that are very unlikely to complete, and the time saved goes to a cheaper rung that can.
Passing the remaining budget as the call’s own timeout closes the loop. A dependency given a fixed timeout can outlive the request deadline, which is how a cancelled request keeps consuming a connection for another ninety seconds.
Validation & Testing
def test_remaining_decreases_and_expires():
d = Deadline.start(0.05)
assert d.remaining_s() > 0
time.sleep(0.06)
assert d.expired()
def test_sub_deadline_never_exceeds_the_parent():
parent = Deadline.start(1.0)
child = parent.sub(0.9, "retrieval")
assert child.budget_s <= parent.remaining_s() + 1e-6
def test_admission_declines_what_cannot_finish():
d = Deadline.start(0.2)
time.sleep(0.15)
assert not admit(d, typical_s=1.0, name="geometry")
def test_expired_parent_refuses_to_create_a_child():
d = Deadline.start(0.01)
time.sleep(0.02)
try:
d.sub(0.5, "retrieval")
except DeadlineExceeded:
return
raise AssertionError("an expired deadline must not yield a child budget")
def test_dependency_receives_the_remaining_time():
seen = {}
def fake(_arg, timeout=None):
seen["timeout"] = timeout
return "ok"
d = Deadline.start(1.0)
call_with_deadline(fake, d, 0.1, "fake", "arg")
assert 0 < seen["timeout"] <= 1.0
The last test is the one that catches the most common regression. Passing a constant timeout to a dependency is the default in most client libraries, and forgetting to override it means the deadline governs the orchestrator while every actual network call ignores it.
Gotchas & Edge Cases
Wall-clock arithmetic. A clock adjustment makes a wall-time deadline expire immediately or never. Use a monotonic source, and be aware that it does not survive process boundaries — deadlines crossing a service boundary must travel as a remaining duration, not as an absolute timestamp.
Deadlines passed as absolute times between machines. Clock skew between hosts turns a shared absolute deadline into a different budget on each. Send the remaining milliseconds and let the receiver start its own clock.
A sub-deadline that outlives its parent. Fractional child budgets are convenient and must be clamped to what remains, or a stage granted “half the budget” late in a request gets more time than the request has left.
Retries inside a deadline that ignore it. A client library retrying three times with its own backoff will happily spend the entire budget without consulting it. Disable library-level retries and do retry explicitly, checking the deadline between attempts.
Cleanup work counted against the budget. Writing a log line or emitting a metric after the deadline expires is correct and should not raise. Check the deadline before starting work, not in finally blocks.
Typical durations that drift. Admission control depends on knowing what a call usually costs, and that number changes. Derive it from observed latency percentiles rather than hard-coding it, and refresh it periodically.
Frequently Asked Questions
What should the request budget be?
Whatever a user will actually wait for, which for an interactive agent is a few seconds and for a batch job may be minutes. Deriving it from dependency latencies is backwards: the budget is a product decision, and the dependencies then have to fit inside it or be replaced by cheaper rungs. A budget set by adding up what the current implementation happens to take is not a budget, it is a description.
Should every stage get a fractional sub-deadline?
Only where a stage must be prevented from consuming everything — typically the first, expensive stage of a ladder. Elsewhere, passing the parent deadline directly is simpler and lets a fast stage return its unused time to the ones after it. Fractional splits allocate optimistically and, when an early stage finishes quickly, leave later stages artificially constrained.
How does this interact with server-side cancellation?
It should trigger it. Passing the remaining budget as the call's timeout lets the client abandon the call, but the server may keep working unless it is told; where the protocol supports a deadline header or cancellation token, propagate it. Otherwise an abandoned request continues to consume database connections and geometry workers on behalf of a user who has already been answered.
What should be reported when the deadline is the reason for a refusal?
The fact and the stage, not the internal numbers. "This took longer than the time available; the boundary check did not complete" is actionable — the user can retry or narrow their question. Reporting the budget and the elapsed milliseconds exposes implementation detail while answering none of the questions a reader has.
Should the deadline cover work after the answer is produced?
No. Logging, metrics, cache writes and audit records happen after the user has been served and must not be cancelled by an expired deadline — nor should they extend it. Structure the request so the deadline governs everything up to producing the answer, and treat post-answer work as a separate concern with its own, generous limits. Mixing them produces the worst outcome available: an answer computed successfully and then lost because the audit write ran out of time.
How should a budget be divided when stages run concurrently?
They share the same deadline rather than dividing it, since concurrent stages are spending wall-clock time together rather than in sequence. What does need attention is the aggregate: three concurrent calls each given the full remaining budget can all still be running when it expires, so the coordinating code must cancel the stragglers rather than waiting for them. A deadline that is checked only before starting work leaves that gap wide open.
Related
- Up to the parent topic: Fallback Routing for Geospatial Queries
- Implementing Fallback Routing for Failed Spatial Queries
- Related topic: Cost and Latency Budgets for Spatial Agents
- Related topic: Error Mapping for Spatial API Calls