A polygon boundary is an ordered ring of coordinates, and every decision about how that ring becomes text changes both its cost and how reliably a model can work with it. This guide covers the four decisions that matter — where the ring starts, which vertices survive, how the numbers are written, and what happens when the result does not fit — as the working core of geometry tokenization strategies.
When to Use This Approach
Apply it whenever a boundary reaches a prompt. If the model only ever needs to know which region something is in, cells or a name are cheaper and better; boundary tokenization is for questions where the shape itself is the subject.
| The model needs to | Send | Because |
|---|---|---|
| Know which region a point is in | A name or a cell identifier | The boundary is irrelevant to the answer |
| Compare two shapes | Both boundaries, canonical form | Ordering differences read as shape differences |
| Describe a shape’s character | A reduced boundary | Vertex count far exceeds what the description needs |
| Reproduce or edit a shape | A structured exact form | Compact text breaks under editing |
| Measure something | Nothing — compute it | A model measuring from vertices is guessing |
The last row is worth stating plainly because it is the most common misuse. A model handed a boundary and asked for its area will produce a number, and the number will be wrong in a way that scales with the shape’s complexity.
Implementation
The tokenizer canonicalises, reduces, formats and budgets, in that order, and reports what it did at each step.
import logging
from dataclasses import dataclass
from shapely.geometry.polygon import orient
from shapely.errors import GEOSException
from shapely.validation import make_valid
log = logging.getLogger("boundary_tokenizer")
@dataclass(frozen=True)
class Boundary:
text: str
vertices: int
original_vertices: int
places: int
note: str
def canonical_ring(geom):
"""Fixed exterior direction and a deterministic starting vertex."""
try:
oriented = orient(geom, sign=1.0) # exterior counter-clockwise
except Exception as exc:
log.info("could not orient geometry (%s); using it as given", exc)
oriented = geom
coords = list(oriented.exterior.coords)[:-1] # drop the repeated closing vertex
if not coords:
return []
start = min(range(len(coords)), key=lambda i: (coords[i][0], coords[i][1]))
return coords[start:] + coords[:start] + [coords[start]]
def reduce_vertices(geom, target: int, to_metric, from_metric):
"""Simplify toward a vertex target by searching tolerance, preserving topology."""
if len(geom.exterior.coords) <= target:
return geom, 0.0
projected = to_metric(geom)
lo, hi = 0.0, max(projected.bounds[2] - projected.bounds[0],
projected.bounds[3] - projected.bounds[1]) / 20.0
best, best_tol = projected, 0.0
for _ in range(12): # bisection: 12 steps is plenty
mid = (lo + hi) / 2
try:
candidate = projected.simplify(mid, preserve_topology=True)
except GEOSException:
break
if candidate.is_empty:
hi = mid
continue
if len(candidate.exterior.coords) > target:
lo = mid
else:
best, best_tol, hi = candidate, mid, mid
if not best.is_valid:
best = make_valid(best)
return from_metric(best), best_tol
Bisecting on tolerance rather than picking one is what makes the vertex target meaningful. A fixed tolerance produces wildly different vertex counts across features of different sizes, whereas a target vertex count produces a boundary whose cost is predictable — which is what a budget needs.
The formatting step writes the ring as pairs and enforces the budget by reducing further, never by cutting the string.
def tokenize_boundary(geom, places: int, budget: int, count_tokens,
to_metric, from_metric, target: int = 64) -> Boundary:
"""Canonical, reduced, budgeted. Never truncates a coordinate list."""
original = len(geom.exterior.coords)
reduced, tol = reduce_vertices(geom, target, to_metric, from_metric)
ring = canonical_ring(reduced)
def render(r, p) -> str:
pairs = ", ".join(
f"{round(x, p):.{p}f}".rstrip('0').rstrip('.') + " " +
f"{round(y, p):.{p}f}".rstrip('0').rstrip('.')
for x, y in r)
return f"POLYGON(({pairs}))"
text = render(ring, places)
if count_tokens(text) <= budget:
note = "" if tol == 0 else f"simplified at {tol:.2f} m"
return Boundary(text, len(ring), original, places, note)
for fewer_places in range(places - 1, 1, -1): # cheapest reduction first
text = render(ring, fewer_places)
if count_tokens(text) <= budget:
return Boundary(text, len(ring), original, fewer_places,
f"precision reduced to {fewer_places} places")
for fewer_vertices in (32, 16, 8):
reduced, tol = reduce_vertices(geom, fewer_vertices, to_metric, from_metric)
ring = canonical_ring(reduced)
text = render(ring, max(2, places - 2))
if count_tokens(text) <= budget:
return Boundary(text, len(ring), original, max(2, places - 2),
f"reduced to {len(ring)} vertices at {tol:.2f} m")
minx, miny, maxx, maxy = geom.bounds
log.info("boundary reduced to its extent to fit %d tokens", budget)
return Boundary(f"BBOX({minx:.4f} {miny:.4f}, {maxx:.4f} {maxy:.4f})",
4, original, 4, "replaced by its bounding extent")
Reducing precision before reducing vertices is the right order because precision loss is invisible at the scales most questions care about, while vertex loss changes the shape’s outline. Both are reported, so a consumer can tell which happened.
Validation & Testing
from shapely.geometry import Polygon
def test_canonical_form_is_rotation_invariant():
a = Polygon([(0, 0), (0, 1), (1, 1), (1, 0)])
b = Polygon([(1, 1), (1, 0), (0, 0), (0, 1)])
assert canonical_ring(a) == canonical_ring(b)
def test_budget_is_met_without_truncation():
out = tokenize_boundary(COMPLEX_POLYGON, 5, 200, count_tokens, to_m, from_m)
assert count_tokens(out.text) <= 200
assert out.text.count("(") == out.text.count(")")
def test_reduction_is_reported():
out = tokenize_boundary(COMPLEX_POLYGON, 5, 200, count_tokens, to_m, from_m)
assert out.note and out.vertices < out.original_vertices
def test_vertex_target_is_approximately_met():
reduced, _ = reduce_vertices(COMPLEX_POLYGON, 64, to_m, from_m)
assert 32 <= len(reduced.exterior.coords) <= 80
The first test is the one worth writing first: rotation invariance is easy to lose to a refactor, and its absence shows up not as an error but as a cache hit rate that quietly falls to zero.
Gotchas & Edge Cases
Interior rings dropped silently. A polygon with holes has more than one ring, and a tokenizer that renders only the exterior produces a shape that is wrong in exactly the region a hole was meant to exclude. Render all rings, and if the budget forces a choice, drop the smallest holes and say so.
Canonical start chosen by index rather than by geometry. Starting at the first vertex in the file makes the serialisation depend on the exporter. Choosing the extreme vertex, as above, makes it depend only on the shape.
Simplification applied in degrees. A tolerance in degrees is a different distance at every latitude, so a corpus spanning a continent is reduced unevenly by a constant that looks uniform. Project first.
A vertex target smaller than the shape needs. Reducing a complex coastline to eight vertices produces a triangle that is technically a polygon and describes nothing. Floor the target, and prefer the extent rung over an absurd reduction — an extent at least announces itself.
Rings that are not closed. Some exporters omit the repeated closing vertex and some readers require it, so a tokenizer that assumes one convention will occasionally emit a ring no parser accepts. Close explicitly when rendering, as the canonical form does.
Multipart geometries flattened to their largest part. Common, convenient, and wrong for anything that spans islands or detached parcels. Render the parts that fit, count the ones dropped, and report both.
Frequently Asked Questions
Is delta encoding worth it for long rings?
Sometimes, and less often than it looks. Writing each vertex as an offset from the previous one shortens the numbers considerably for a dense ring, which does reduce tokens — but it produces text no standard parser accepts and that a model handles less reliably, since it must reconstruct absolute positions to reason about them. Reserve it for cases where the ring is the whole payload and a custom parser already exists on the other side.
Should the boundary or the centroid be sent for a "where is it" question?
Neither, usually — send the name and let the geometry stay in the database. Where the position genuinely matters, a representative point is cheaper and more robust than a boundary, and unlike a centroid it is guaranteed to lie inside the shape, which matters for concave features where the centroid falls outside.
How many vertices does a model actually need?
Far fewer than most corpora carry. For describing character and arrangement, a few dozen is ample; beyond that the additional vertices mostly encode survey precision that the question never touches. The exception is comparison — two shapes reduced to different vertex counts compare badly — which is another reason to use one target across the corpus rather than per feature.
What should happen when a boundary crosses the antimeridian?
Split it before tokenizing. A ring whose longitudes jump from 179 to −179 is rendered as an enormous shape spanning the globe by every naive reader, including the model. Splitting produces two rings that are each unremarkable, and the note should record that the feature was split so nothing downstream treats them as separate features.
Should the vertex count reach the prompt alongside the boundary?
Yes, when the boundary was reduced, and it costs almost nothing to include. A model told that a shape has been reduced from three hundred vertices to sixty-four will hedge appropriately about fine detail; one handed sixty-four vertices with no context treats them as the whole truth. The same applies to the tolerance: a stated figure in metres is far more useful to a reader than a general note that simplification occurred.
Related
- Up to the parent topic: Geometry Tokenization Strategies
- Coordinate Precision Versus Token Cost
- Well-Known Text, Structured Objects and Cell Identifiers
- Related topic: Context-Window Optimization for Maps