Three representations dominate geometry in prompts, and they differ by roughly a factor of six in cost and completely in what they preserve. Choosing between them is not a matter of taste: each one makes a different set of questions cheap and a different set impossible. This guide compares them on the axes that matter for a prompt pipeline, as part of geometry tokenization strategies.
When to Use This Approach
Choose once, for the corpus, and convert at ingestion. Mixing representations means every parser and every prompt must handle both, and the model spends capacity distinguishing forms rather than reasoning about places.
| Property | Compact text | Structured object | Cell identifiers |
|---|---|---|---|
| Token cost for one polygon | Low | High — about 1.6× | Lowest — about 0.4× |
| Exactness | Exact | Exact | Lossy by construction |
| Model manipulation reliability | Moderate | High | High for set operations |
| Containment and adjacency | Needs a predicate | Needs a predicate | Native, cheap |
| Boundary questions | Answerable | Answerable | Not answerable |
| Human readability in a log | Good | Verbose | Opaque |
The row that decides most cases is the second-to-last. Cell identifiers answer “is this inside that” and “what is next to this” almost for free and cannot say anything about where a boundary actually runs, because the boundary has been replaced by a staircase.
Implementation
The conversion functions below are the ones a corpus needs at ingestion. Each returns the text plus a note about what was lost, so the prompt layer can state it.
import logging
from dataclasses import dataclass
log = logging.getLogger("geometry_representation")
@dataclass(frozen=True)
class Rendered:
text: str
form: str # compact | structured | cells
lossy: bool
note: str
def as_compact_text(geom, places: int) -> Rendered:
"""Exact geometry, minimal syntax. The default for most corpora."""
def fmt(v: float) -> str:
return f"{round(v, places):.{places}f}".rstrip("0").rstrip(".") or "0"
rings = [", ".join(f"{fmt(x)} {fmt(y)}" for x, y in ring) for ring in _rings(geom)]
return Rendered(f"POLYGON(({'), ('.join(rings)}))", "compact", False, "")
def as_structured(geom, places: int) -> Rendered:
"""Exact geometry in an object form a model can edit without breaking syntax."""
import json
coords = [[[round(x, places), round(y, places)] for x, y in ring]
for ring in _rings(geom)]
body = {"type": "Polygon", "coordinates": coords}
return Rendered(json.dumps(body, separators=(",", ":")), "structured", False, "")
def as_cells(geom, resolution: int, to_cells, max_cells: int = 64) -> Rendered:
"""A cell cover of the geometry. Cheap, approximate, and explicitly labelled."""
try:
cells = list(to_cells(geom, resolution))
except Exception as exc: # a library failure must not lose the feature
log.warning("cell cover failed at resolution %d: %s", resolution, exc)
return as_compact_text(geom, 5)
while len(cells) > max_cells and resolution > 1:
resolution -= 1
try:
cells = list(to_cells(geom, resolution))
except Exception:
break
note = (f"approximated by {len(cells)} cells at resolution {resolution}; "
"boundary detail is not represented")
return Rendered(" ".join(cells), "cells", True, note)
Falling back to exact text when the cell library fails is the right direction: an exact representation is never wrong, only expensive, whereas dropping the feature or emitting an empty cover is both.
The coarsening loop deserves a note. A cell cover of a large or thin polygon can run to thousands of identifiers, which defeats the point of using cells at all; stepping the resolution down until the cover fits keeps the cheap representation cheap, at the cost of a coarser approximation that the note reports.
Validation & Testing
def test_compact_and_structured_agree_geometrically():
from shapely import wkt
import json
from shapely.geometry import shape
a = wkt.loads(as_compact_text(POLYGON, 5).text)
b = shape(json.loads(as_structured(POLYGON, 5).text))
assert a.equals_exact(b, 1e-5)
def test_cell_form_is_labelled_lossy():
r = as_cells(POLYGON, 9, to_cells)
assert r.lossy and "boundary detail is not represented" in r.note
def test_cell_cover_coarsens_rather_than_exploding():
r = as_cells(LARGE_THIN_POLYGON, 12, to_cells, max_cells=64)
assert r.form == "compact" or len(r.text.split()) <= 64
def test_library_failure_falls_back_to_exact():
def broken(_geom, _res):
raise RuntimeError("cell library unavailable")
r = as_cells(POLYGON, 9, broken)
assert r.form == "compact" and not r.lossy
The first test is the one that catches a whole class of conversion bugs, because it compares two independent renderings of the same geometry rather than checking either against a fixture string. Fixture strings encode formatting decisions and break whenever formatting changes; a geometric comparison breaks only when the geometry does.
Gotchas & Edge Cases
Structured output that a model has edited. The main reason to pay 1.6× is that a model asked to modify geometry breaks compact text far more often than it breaks an object form. If nothing in the pipeline asks the model to produce or edit geometry, that advantage is not being used and the extra cost is pure.
Cells used for a boundary question. The failure is silent: the model answers confidently from a staircase. Where cells are used, restrict them to membership and adjacency, and route boundary questions to exact geometry.
Ring orientation differing between forms. The two exact forms have different conventions in common use, and a corpus that converts between them without canonicalising will produce two token sequences for one shape. Fix orientation at ingestion.
Compact text with excess whitespace. Some writers emit a space after every comma and around every parenthesis, which is a measurable fraction of the tokens in a coordinate-dense string. Emit a canonical minimal form rather than whatever a library defaults to.
Cell resolution chosen globally. One resolution across a corpus of mixed feature sizes produces a four-cell cover for a parcel and a four-thousand-cell cover for a county. Choose resolution from the feature’s extent, cap the cell count, and report both.
Frequently Asked Questions
Can more than one representation be stored?
Stored, yes; sent to the model, no. Keeping exact geometry in the database and a cell index alongside it is a normal and useful design — the cells serve filtering and the geometry serves measurement. What causes trouble is putting both in a prompt, where the model must decide which one to trust and will occasionally reason over the approximation while quoting the exact form.
Which form should be used when the model must return geometry?
The structured object form, almost always. A model closing three levels of parentheses correctly in a long coordinate list is doing something it is not reliably good at, and the failure mode is unparseable output. An object form has redundancy that makes errors detectable and often repairable, and the token premium is small relative to the cost of a failed generation.
Do cell identifiers help with retrieval as well as prompts?
Substantially, and that is often the better use of them. A cell identifier is an exact-match token, which means the lexical half of a hybrid retrieval system can find every document about a region with a keyword lookup rather than a spatial join. That is a real capability, and it does not require the cells to appear in any prompt.
How much does the choice matter compared to precision?
Comparable, and they compound. Moving from a structured object at seven decimal places to compact text at five is roughly a threefold reduction, which is larger than either change alone. Make both decisions at the same time and measure the combined result on real features, since the interaction between representation overhead and coordinate length is not obvious from either number.
What should a corpus do about geometry types other than polygons?
Apply the same choice consistently. Points are cheap in every form and the decision hardly matters; lines behave like polygon rings and inherit the same trade-offs; multipart geometries multiply the representation overhead and are where the structured form’s verbosity hurts most. If one type dominates the corpus, let it drive the choice, and convert the minority types into the same form rather than special-casing them.
How should the chosen form be recorded?
In the chunk metadata, alongside the precision and any simplification tolerance. A reader — human or machine — encountering geometry in a retrieved chunk needs to know whether it is exact, and that fact belongs with the data rather than in a pipeline document. It also makes a later migration tractable: converting a corpus is straightforward when every record says what it currently is.
One practical note on migration. Converting a corpus between exact forms is mechanical and safe; converting to or from cells is not, because the cell form cannot reconstruct what it approximated. If cells are stored as the only representation, that decision is irreversible for the affected features, which is a reason to keep exact geometry in the database even when cells are what reaches the prompt.
Whichever form is chosen, write the conversion in one place and call it from everywhere. The failure this avoids is not conversion errors but drift: two call sites that format geometry slightly differently produce a corpus with two dialects of the same representation, which is all the cost of mixing forms with none of the benefit.
Related
- Up to the parent topic: Geometry Tokenization Strategies
- Coordinate Precision Versus Token Cost
- How to Tokenize Polygon Boundaries for Transformer Models
- Related topic: Hybrid Spatial and Keyword Retrieval