The generator’s model of record is a relational fact base. The store’s DDL (graphitron-model.sql in graphitron-model) is the model’s what: which relations exist, their keys, their constraints, and the comment on every one of them, rendered browsable as the generated schema reference, one page per relation-name family, generated at build from those same comments. This page is the why: the discipline that shapes those relations and the invariants that keep the model debt-free while the transitional classification walk drains. Where this page and the DDL disagree, the DDL wins. What runs when is the Pipeline overview; this page owns the rules a model change must obey. Every claim below names its enforcer, the live test or gate that fails when the claim breaks; a rule without one is not on this page. A gentler first contact with the naming discipline, examples and metaphors included, is Naming the row.

The discipline has an executable form

The facts are typed, keyed relations in an embedded H2 database booted from the DDL, read and written through jOOQ classes generated from that same DDL. That closes the usual objection to a SQL layer, that it makes the model stringly-typed and moves exhaustiveness to runtime: the generated classes are the containment, graphitron-model builds before the generator core, and a DDL edit fails javac in every consumer that touched the changed relation. There are no schema migrations and no persisted state of record; a warm store is reconciled per source hash and an incompatible upgrade opens a fresh file, so compile time is the only compatibility surface the schema has, and changing the model is editing the DDL and following the compiler.

Referential integrity splits by what it can enforce. Capture-structural edges, where the walk writes the child while standing on the parent (a field row’s declaration site, an application’s host directive), are declared `FOREIGN KEY`s. A reference the author spells by name deliberately carries no FK and is a detection query instead, because a dangling author reference is a diagnostic to report, never a capture bug to crash on.

Enforced by: the compiler over the generated jOOQ surface; FactSchemaGateTest holds the store’s own invariants (total comment coverage, dense ordinals, verbatim-transcription twins for every decode, the graph-partition rules), so a failure there is a capture bug or a DDL defect, never an author error.

The key discipline: natural keys, rendered coordinates

Base relations key schema elements by the spec grammar’s own Name columns, per relation and all non-null: graphql_field keys (graph_name, type_name, field_name), and its kin follow the same shape. There is no surrogate id anywhere in the DDL; where a tie-breaker is needed it is an ordinal inside the natural key, never decoration on a counter. The keys are what the model joins and scans on: every field of a type is a prefix scan on the type name, and every graph-partitioned key leads with graph_name.

The canonical coordinate string (Foo, Foo.bar) appears in the store only as a rendering of those columns. The stored coordinate column on the diagnostic read surface says so in its own comment: a rendered value computed from the stored columns, never the key. A rendered key cannot drift from the columns it names, which is what makes it safe as a stable id to hand outward.

Authored facts are definition-keyed; derived bindings are use-keyed. An input field’s authored facts live at its own member coordinate, the only key available where the author’s cursor sits, because one input type is consumed by many fields. The use-site resolution, which table an input binds against and through which occurrence path, depends on the consuming field, so it is a derived join over definition and consumer, shipped as the materialized intent_input_occurrence_path pair rather than as a dotted-path key.

Enforced by: the PRIMARY KEY`s themselves, at capture; `FactSchemaGateTest.everyRelationLeadsWithItsPartitionDimension, everyGraphKeyedRelationReachesTheAnchor and aRunWritesOnlyUnderItsOwnGraph for the partition half; FactSchemaGateTest.commentCoverageIsTotal keeps the rendered-not-stored statements written on the columns that carry them.

Facts, not leaves

Each fact is an independent functional dependency of its coordinate, found by its own walk: where the source object arrives, what the field returns, which operations it triggers, whether the value crosses a table. A capability is added by adding a fact relation, never a new leaf type, because facts add where leaf types multiply: welding independent axes into one classified leaf produces the cross-product of the axes, and the leaf zoo is exactly the denormalized view that multiplication built (repeating groups for multi-column values; duties welded onto a leaf that functionally depend on another coordinate’s query). The sealed leaf model the classification walk still produces is the transitional producer surface of the strangler migration, drained rather than extended; a new fact lands in the store.

Enforced by: FactCaptureAgreementTest.everyRelationIsRegistered, whose driver enumerates every generated relation and fails on one with no registered agreement source, with a reverse leg failing on a registration the DDL no longer declares, so a fact relation can neither arrive unchecked nor linger as a phantom.

Name the row, not the question

A relation named for a consumer’s question inherits that question’s shape, and the damage is not visible from inside the work. Its grain becomes whatever the caller’s return value was. Its derivation absorbs every special case the caller had, arriving as one relation with a procedure behind it rather than a fact with a grain. And because no independent source states the answer, the only available oracle is the code being replaced, so a migration’s own predecessor becomes normative and whatever bugs it has become pinned invariants. The three symptoms travel together, and each one looks like a local modelling choice on its own.

The check runs before the DDL and costs one sentence: say what a single row asserts, without naming a consumer, a generator pass, or an existing class. "This classfile declares this supertype through this clause" passes. "This class offers this slot under this name" passes. "The class the resolver would bind to this type" fails, and it fails while the fix is still free. The sentence has to survive at the column grain as well, and a column encoding a set is where it stops surviving: the row then asserts something about several things at once, so nothing said about one subject finishes it honestly, which is the rule under "Derived reads are views, not stored facts" read one grain down.

What survives the check is usually several facts where the question suggested one, each anchorable against its own source and composable by the next consumer, which wanted a different combination. The relation the question asked for often survives too, radically thinner, once the clauses that were really separate observations have moved out beside it. Watch for the ones that were never facts at all: a filter one caller applies, and an agreement between two facts, which is a detection rather than a step.

The corollary is about oracles. Fidelity to a predecessor is evidence, not a specification. What that forbids is the total-agreement test: asserting that the replacement equals the predecessor installs the old code as the standard and pins whatever bugs it has as invariants. What it asks for instead is named fixtures, agreement asserted where the two are meant to agree, and each deliberate departure asserted in the direction it was meant to go, so a difference stays adjudicated rather than outlawed.

Where the predecessor’s answer lives is a second question, and a smaller one than it looks. Writing it into the store buys a comparison that runs over any corpus a run touches and answers while the replacement is half built; that is worth its capture cost only where something actually diffs the pair at a scale no test holds in hand. Nothing has, so far: the one shipped case stored its predecessor’s answer on every capture, was diffed only by four hand-written fixtures, and retired with the comparison moved into that test’s own JVM and its shape unchanged. Take the store-side form when a corpus-scale diff is a thing you will run, and the in-memory form otherwise; the shape above is what neither form may give up.

Not mechanically enforced: the check is a sentence said out loud before the DDL exists. The gates catch downstream shapes (an unregistered relation, a derivation with no anchor, a stored value that should be a view), never the naming that produced them.

Provenance: every source is a fact, and the resolved value is another

Do not assume a fact’s origin is a column on the fact, and do not assume that two sources of "the same" value are one fact at all. They are at least three. What the author wrote is one fact, stated where the author wrote it. What another corpus publishes is a second, stated where that corpus was read. What the generator should use is a third, derived from the others, and it is the only one a consumer of the resolved answer reads. A third source is a fourth fact and changes none of this. Collapsing them into one relation with an origin column loses every source but the winner, because a row can then hold only the value that won. Keeping the sources and calling their union the answer loses the derived fact, because a union is not a decision.

The derived fact owes an answer where its sources disagree, and that answer is the substance of the derivation rather than a detail of it. Precedence is one answer, and node identity takes it: an author writing @node(typeId:) beats the same table’s published constant, because the author is saying something about this graph where the catalog is saying something about a database that many graphs may read. Refusal is another answer, and so is a row naming the contradiction so somebody can be told. A rule that picks silently where its sources conflict has taken the decision without recording that there was one, and the next reader cannot tell a resolved disagreement from an absence of one.

Record which source answered, and record it on the derived fact. It reads like a label and is not one: it is the way back. A reader holding type_id_origin = JOOQ_METADATA knows to join sql_node_metadata for that table, where the rest of what the generated class published is waiting, including the constants this rule did not need; a reader holding SDL_DECLARED knows to reach the entry instead, where the written position is, and can point at the author’s own text. Without the column both are re-derivations of a decision the rule already took, and each reader takes it again and may take it differently. This is the answer to the objection that a tag no consumer branches on is inventory: consumers do not branch on it, they follow it, and a fact that opens a door is worth its column even where nothing behind that door is wanted yet.

How the derivation is realised is a separate question, and the storage rule below answers it rather than this section; The modeling discipline walks the same ground from provenance, for a contributor about to add a relation. A view states the rule where nothing forces otherwise. A stored relation is right where the derivation establishes a grain something else keys into, or where the cost of re-deriving it was measured rather than assumed. Nodehood is the shipped case of the stored form: graphitron_node and graphitron_node_keycolumn merge the authored and the published populations into the relations that are authoritative for nodehood, each carrying which tier answered, because that merge is what every consumer reads and because a tiered union re-expanded once per probe is the cost a stored derivation exists to remove. The sources survive it: graphitron_node_entry still holds what the author wrote, null where they wrote nothing, so a surface reporting on the author’s own text reads that relation and not the merge.

The claim stratum is the shipped case of the other shape, where the derivation stays a view: intent_authored_field_claim and intent_authored_type_claim union one arm per claiming graphitron_ relation, intent_column_match_claim carries the structural reading with its join witnesses as payload, intent_resolved_field_claim is the reduction, and intent_authored_claim_conflict detects coordinates carrying contradictory claims. That relation is total over the authored claims and carries no population filter: an authored contradiction is a contradiction wherever it sits, so each consumer joins the population its own question needs (the build-error surface joins the classification domain, because only the emitted surface can fail a build; the editor’s diagnostic arm reads the rows ungated). The conflict reduction lives in the view’s SQL; AuthoredClaimConflicts mints the typed rejection from the closed verdict vocabulary and the coordinate’s own claim rows, and applies the build-error population. The message is minted rather than derived because its naming order is the claim enum’s declaration order, which is not a captured fact of any graph and so is no view’s to express; a capture-cadence writer stores the same rejection in intent_authored_claim_rejection for the diagnostics surface to read as a plain column, which is why that surface’s claim-conflict arm looks like its captured siblings instead of assembling a sentence in SQL. The one-slot case is graphitron_node_entry: type_id holds what the author wrote and stays null where they wrote nothing, so a surface reading the relation sees authored data only, by design, and the resolved value falls back to the type-name rule the relation’s own comment names as a derivation.

The consumer split is which population each reads: code generation wants the resolved answer, the editor gives feedback only on what the author wrote, the knowledge surface cites the authored source and reports the resolved value.

Description looks like the case for a column, and it splits. Where the walk that produces the row reads the description in the same pass (the SDL docstring, the catalog COMMENT), the description is co-sourced with the entity and belongs on the entity’s own relation. Javadoc is the counter-case, and it was once listed here as though it were the same thing: it comes from a parse of the consumer’s sources, not from the bytecode scan that produces code_class, and it moves when a file is saved rather than when a build runs. Two walks on two cadences, so its own relation in the java_ family, joined by name, and never a column on the census. The discriminator is not whether the two values describe one subject; it is whether one walk read them both. The same question, "where did this come from", gets opposite shapes depending on whether the origins are independent walks; the instinct must be "pick", not "always a column".

Enforced by: FactCaptureAgreementTest registers the whole intent_ stratum under its derived arm, with per-view anchors pinning each derivation against hand-written expectations the view cannot produce by construction. An anchor sits at the layer whose question it asks. Agreement with the transitional walk stays beside the walk (AuthoredClaimConflictsTest, ColumnMatchShadowTest, DemandShadowTest, InputOccurrenceShadowTest in rewrite/derive) and retires with it; what a view returns given rows is pinned in graphitron-model, the module whose DDL declares it, against a store seeded row by row (ColumnMatchClaimTest, AuthoredClaimTest, DemandRuleTest, InputOccurrenceOverrideTest, SeparateFetchRuleTest, TypeBackingSeedTest, TypeBackingTest and their siblings in model/intent). A view reading a fact about the outside world keeps a third anchor beside the crawler that produces it, asserting that real inputs reach the relation in the shape the rule reads.

The three strata: capture, derivation, queries

The store has three strata. Capture transcribes facts from a corpus. Derivation computes further facts from captured ones. Queries read facts to serve a goal. Which one a row belongs to is decided mechanically, by a single test: a row that can be recomputed from captured facts alone is a derived fact and must not be captured. Stated once, that test settles the cases the next section argues one at a time, and it settles the ones no section has reached yet.

The numbered form is this axis and only this axis. "Stratum one", "stratum two" and "stratum three" mean the three above; an unnumbered "the X stratum" means something else, a section of the DDL or a layer inside one family. Without that rule a reader meeting "the diagnostics stratum" a few paragraphs from "stratum one" counts four. This page carries four unnumbered uses, and re-reading them under the rule is the whole of what the numbering does to them: "the claim stratum" and "the whole intent_ stratum" name a family, and "the diagnostics stratum" (twice) names an arm set that spans three of the buckets below. A fifth use, "the derived stratum", was stratum two under an older name, so the numbering replaced it outright instead of re-reading it. Take that set from the file rather than from this sentence; a count stated in the section that argues against unguarded counts is the wrong thing to trust.

A relation’s stratum is decided by what its rows are a function of, never by what program computes them. A materialization whose producer is a Java walk is stratum two if its inputs are captured facts, and materialization is sanctioned above where a view cannot serve. Saying this in the same breath as the assignment matters, because "the decoded reading is a derivation" otherwise reads as a demand that a thousand lines of decoding become SQL, which is not what the stratum claims.

A rule reading two corpora is stratum two, and this is the one corollary of the recompute test that a gate can close. A crawler answers for a corpus that exists independently of the others: the .graphqls files are what they are whether or not a jOOQ package was generated, and the generated classes are what they are whether or not any schema mentions them. So the rows a crawler writes about its own corpus must not vary with another corpus’s contents, and a rule whose answer needs both is a derivation over their captured facts rather than a decision either crawler may make. Federation’s key synthesis is the worked case: it fires on nodehood, and nodehood conjoins the SDL’s claim (@node, or @table plus implements Node) with node-identity metadata a generated jOOQ class publishes. Run inside the SDL walk it made an unchanged .graphqls file write different graphql_ and graphitron_ rows when a generated class changed, which is stratum two running inside stratum one and landing its output where nothing tells it apart from a transcription.

No foreign key could have caught that, which is why the rule needs a gate of its own rather than a constraint. A key constrains references, and the schema already refuses to model SDL-to-jOOQ resolution as one: there is no graphql_ to sql_ edge anywhere, and a @table(name:) row holds a string a crawler transcribed rather than a pointer at sql_table. A cross-corpus read inside a crawler adds no reference at all. It changes which rows exist, which no constraint expressible in the DDL can reject.

The assembling gatherer reads two corpora, and says so. Its verdict is about the corpus as the generator composes it, and the composition is configured: the inputs' tags and description notes are applied before assembly, so the same documents under a different <schemaInput tag> are a different schema to judge. A crawler’s rows may not vary with another corpus, so meta_gatherer_corpus declares configuration beside sdl on the graphql-assembly roster, which makes the configuration one of that gatherer’s own corpora rather than a corpus it leaks. The decode does not take the same widening: it walks the merged corpus the composition started from, the documents and nothing else, so the sdl gatherer’s roster stays one corpus and the tag configuration cannot reach a transcription by construction. Composing the verdict a second time, reader-neutrally, is what the store did before, and it reported a federation `@link’s imports as undeclared directives the build never sees.

Enforced by: CaptureCorpusIsolationTest captures one registry twice, once with the jOOQ catalog and once without, and requires every graphql_ relation to hold identical rows across the two. Beside them it holds the as-written half of graphitron_, the relations restating one directive application in graphitron’s vocabulary, whose rows are a function of one document and of nothing else. The scope is a property of the rows rather than of who writes them: those rows join nothing and resolve nothing, so catalog-independence is something they have by construction and the gate is what says so. What stays outside is the resolved half, and the roster says why: meta_gatherer_corpus gives the graphitron gatherer no corpus row at all while meta_gatherer_dependency has it depend on both sdl and catalog, so a stage joining an entry against sql_node_metadata answers from captured facts rather than from an input and may reach the catalog legitimately. Holding that half to catalog-independence would forbid the thing it exists to do. Neither half arrives as a list, and the two spellings are why: graphql_ is enumerated off the generated model by prefix and the as-written half by the _entry suffix, so a relation added to either is in scope without being named. What a name cannot do is hold a new decode relation to carrying the suffix in the first place, and that is EntryNamingGuardTest, which reads the decode’s own source and requires every relation it names to be an entry. Note the scope precisely: the gate fires on the cross-corpus subclass of a recompute violation and on nothing else, so a capture-time derivation whose inputs all sit inside one corpus still passes it. And a differential is only as good as its corpus, which is why the widened half runs against a fixture applying every directive the decode writes a relation for: over the transcription’s own fixture most of that half held no row, and a comparison of empty against empty agrees without being able to disagree. The surviving @asConnection expansion is exactly that shape, which is why the inversion the next paragraphs name is a live inversion under a green suite.

Which stratum each family is in

meta_family is the census. It rosters the relation-name families as constant rows, is closed against the observed relations in both directions by the schema gates, and renders one reference page per row; what follows reads that roster rather than restating it, and where the two disagree the roster is right. The buckets:

  • Stratum one, transcription. the declaration relations of graphql_ (the generic directive definitions and their applications at all five locations included, with argument values), sql_, code_, and java_. Each holds what a walk read from one corpus. javac_ is here too, and is the only stratum-one family whose corpus the run itself produced rather than read from the consumer.

  • Stratum two, derivation. intent_, and graphitron_. The graphitron_ relations are decodes of the generic directive applications stratum one transcribes: a @reference path decomposed into steps, a sigil extracted from an argument value, a name: read off @table. Each is a function of captured rows. The macro inversion that used to be the one exception here is corrected. A pre-expansion type expression once sat in the macro’s own provenance relation and in no graphql_ row, which put a stratum-one fact inside a stratum-two family; the transcription now holds what the author wrote and graphitron_field_minted holds the expansion’s replacement, so both are where the recompute test puts them. The federation half was corrected earlier, its rule now a derivation and its provenance relation retired. lint_ is stratum two as well, argued below.

  • Stratum three, queries. No family lands here, and the reason is the roster’s own naming rule rather than a gap. A family is named for whose vocabulary its rows are written in, which makes its rows facts the store owns, which is stratum two. A relation that exists because one consumer wanted one read has no vocabulary of its own to be named for, so it gets no prefix. That is why the two-versus-three boundary is not "a view a consumer reads": intent_resolved_field_claim is a view over captured facts whose own comment calls it what a planning reader eventually joins, and it is stratum two. The prefix-less diagnostic view is the worked case and the roster’s one placement exemption, a read surface unioning arms from several families' vocabularies. Cite it without an arm count; the store’s two statements of that count already disagree, which is an unguarded inventory rather than anything this section introduces. The exemption is over that relation’s name and population and says nothing about what a column on it may hold: coordinate passes the column rule below because its atoms ride the same row, which the view’s own comment argues, and a column truncating a path down to one of its segments failed the same rule on that same relation and was removed.

  • No stratum, scaffolding. rejection_, reifying the transitional classification walk’s verdicts so a derivation can be diffed against them while the migration runs. This bucket is decided before the strata question, and the ordering has to be stated because the strata test would otherwise pull the family into stratum two, since rejection_ holds verdicts the rule below would reach. A relation that exists to be diffed against its own replacement has no stratum whatever its inputs are. That is what scaffolding means here, and the charter of rejection_ already ties its own lifetime to the walk’s retirement. The bucket held a second family until the walk’s backing answers stopped being stored at all; that it can empty is the point of it.

  • No stratum, permanently. store_ and meta_, whose subject is the store itself rather than any corpus or any fact about one. store_ records the run, what it read and what it was built from, and its charter already disclaims transcription in those words; meta_ records the schema. Both need saying, because "no stratum" otherwise reads as "transitional", which is what the pair above is and this pair is not. The meta_ views are also why stratum two has to be stated as a derivation over captured facts rather than as a derivation: meta_relation_family is a derivation over the schema’s own catalog rather than over anything a walk read, so it sits here. It is not stratum three either, meta_ being a family with a vocabulary of its own under the rule above.

A verdict is a conclusion, and the recompute test was stated over declarations, so verdicts need the test extended rather than applied as-is: a transcribed verdict is stratum one exactly while the store does not hold the inputs the verdict was computed from, and stratum two once it does. The rule reaches the verdict residents and does not land them together. graphql_schema_problem splits along the stage column it carries, and not the four ways the column goes. A PARSE refusal is stratum one: its input is a document that has not parsed, and there is no transcription of an unparseable file to recompute a syntax error from. ASSEMBLY is stratum one as well, and it got there by the same move REGISTRY made in the other direction. Its checks are predicates over a registry (that a named type resolves, that an object satisfies the interfaces it claims, that a directive sits where its definition permits, that the schema has a query root), and while the registry it judged was the corpus as written every one of them was a predicate over captured rows. The registry it judges now is the corpus composed with the loading rewrites the generator runs, which includes the definitions the federation library injects for a @link’s imports, and the store does not capture those. `REWRITE, the loading rewrites' own stage, splits by variant. MultipleFederationLinks and TagNotImported are stratum two, recomputable from the @link application entries and store_graph_schema_input.tag, and so is SynthesisedLinkRefused, the registry re-checking operation types the schema extension entries already state. UnsupportedFederationVersion, UnsupportedLinkImport, UnsupportedRename, DeclarationCollision and LibraryDuplicateDeclaration are stratum one, because which definitions a @link imports and which spec versions exist is the library’s knowledge. SourceInTwoInputs is stratum one too: store_graph_schema_input holds one row per recipe entry rather than per match, so the overlap is not recomputable without expanding the globs again. The stratum-two arms are transcribed anyway, on the precedent REGISTRY sets in the same relation: the rewrites have to run to assemble the schema at all and they refuse while running, so a view would state a second time the verdict that actually stopped the composition. REGISTRY is stratum two. It joined that stratum when the entry relations arrived, and how it did is worth following: a registry refusal names a second base declaration whose loser the merge reports without ever offering it to the walk, so while the walk was the only transcription the store did not hold the input. An entry relation is a bag keyed by the position a node was written at, so it holds both declarations, and the collision is recomputable from the pair. That arm changed stratum without anybody moving a row, which is the rule doing the work it exists for: the assignment is a function of what the store holds, so capturing a new input can settle it differently. lint_ is stratum two, and this is the assignment that paid for the rule, because the page said so before it was true. A lint finding was produced by one traversal over the parsed AST that graphql_ transcribes, which made it recomputable from captured facts while it was still stored as a table, and the rule read that as an inversion rather than a description. It has since been acted on: the rules are arms of lint_violation, a view over the entry relations, and the traversal is gone. What the rule predicted from the shape of the inputs is now what the schema says. build_warning_ is the one family the rule does not settle at family grain: the arm is defined by carrying no rule, so its residents share a channel rather than an input set, and whether a given advisory is a function of captured facts is a per-producer question. Stratum one until a producer’s inputs are shown to be captured, decided per resident. A disclosed gap at one of thirteen is the honest form; the failure mode is a clean-looking answer a reader cannot check.

One of those results is still an inversion rather than a description, and naming it is not the same as acting on it: the macro case stores a derivation where the captured fact belongs. Saying which assignment is right is what this page owes; moving the rows is a schema change. The lint_ inversion was the other, and it is worth recording that the rule found it before anybody was looking for it: the family derived what it stored, the rule said so, and the family is now a view.

That schema change has its first instance, so the sentence above is no longer only a promise. Federation’s key synthesis used to run inside the capture walk and land a synthesized @key in the type site’s application relation, its decode in graphitron_federation_key_entry, and a provenance row in a relation whose only purpose was to say which of those rows a macro put there. The rule is now graphitron_synthesized_federation_key, a derivation whose own presence is its provenance, the provenance relation is gone, and stratum one at that coordinate is pure transcription of the SDL. What the instance demonstrates is the cost profile rather than the difficulty: the readers moved, the anchors moved, and no consumer of generated code noticed, the emitted schema’s synthesized @key having always come from the pipeline’s registry rewrite rather than from the store.

The corollary worth stating because it looks like a counter-argument: the foreign keys from graphitron_ into graphql_ are a derivation’s edges to its inputs. They are correct and permanent, and they argue neither for nor against separating the two families. Stated without a count, since an unguarded census rots silently; meta_relation_family and the roster gates are where the enumerable answer lives.

This section makes three claims of different enforceability, and one enforcer line covering all of them would overstate two.

Enforced by: for the claim that a stratum-two decode may not displace the stratum-one transcription, FactSchemaGateTest.theDecodeDoesNotReplaceTheTranscription and the verbatim-transcription twins already cited above, and the gap is worth stating precisely: the gate covers a decode that adds rows beside the transcription, and does not cover a macro that rewrites graphql_field.type_sdl in place, which is exactly the inverted case the macro paragraph below names.

Not mechanically enforced: the recompute test itself, except for its cross-corpus subclass, and the family assignment. One proper subclass of the recompute test is now gated, CaptureCorpusIsolationTest failing on a capture-time rule that reads a second corpus. Everything else the test would reject stays undetected: nothing in the suite fails when a captured relation turns out to be recomputable from facts inside its own corpus, which is why the two inversions named above sit in the store today with every gate green. The list is amended rather than shortened deliberately, since naming which part is covered is what lets a reader tell the two cases apart instead of reading the gate as closing the rule. What would close the rest is making each family’s stratum queryable data on the roster rather than charter prose, the way the roster already closes the family list; that is a schema change and a gate, not a page.

A claim whose enforcer does not close it still belongs here. The preamble’s bar is that a claim names its enforcer, and a disclosed gap satisfies that where silence would not, so a later reader should not read the two paragraphs above as an oversight. This is the one place the page argues with its own preamble rather than obeying it, the preamble asking for the gate "that fails when the claim breaks"; making the disagreement visible is better than quietly satisfying the letter. And the registration arms of FactCaptureAgreementTest are not the stratum’s reflection: they file graphitron_ under the containment arm beside graphql_, because the arms answer how a relation’s contents are pinned, not what its rows are computed from.

The graphql_ family reads a document, not a corpus

Two halves live under this prefix, and they answer different questions. Both are graphql-java’s vocabulary, which is what the prefix names, and neither is the schema API’s: that API is a stricter post-validation view of a schema that assembled, where this family is the document as parsed.

The half that has always been here is keyed by coordinate. graphql_type holds one row per type and graphql_field one row per field, which is the corpus after its documents are merged into one namespace. That shape cannot hold a corpus that fails to merge. Two documents declaring type X is an error a developer commits by saving a file, and the merge reports it by refusing the second definition, so a coordinate-keyed relation has room for one of them and the writer has to choose. Choosing is where a claim, a first-wins rule and an overflow relation came from, and none of them is transcription.

The half beside it is keyed by the position a declaration was written at. Two declarations cannot share one position in one file, so nothing collides, nothing is claimed and nothing is refused: both documents are rows and the collision becomes a query over them. Ordering the candidates is `store_source.mtime’s job, which is the one question a content hash cannot answer and what lets a report name the incumbent and the incursion rather than listing two declarations symmetrically.

The AST entries reference nothing, and their file is named twice

One relation per SDL node kind and site, from graphql_ast_type_declaration_entry through graphql_ast_value_entry. Kind is not always enough to name one: an input value has one shape at all three sites it can be written and a directive application one shape at all five, and each site takes a relation of its own so that a row can name its parent by key. The shared ast segment is the one place a family is subdivided by name, because the graphql_ family now spans two levels: the anchors are schema elements keyed by coordinate, in the specification’s vocabulary, and these are parser nodes keyed by position, in graphql-java’s.

Six things about them are not what a reader would expect.

A node’s parent is a key where the parent is one relation, and a bare position where it is not. A node names the node it was written inside by that node’s position. Both rows come out of one parse of one file and one method writes both, so nothing a document can say breaks the reference, and thirteen of them are foreign keys. Two are not: an input value can be a field’s argument, an input object’s field or a directive definition’s argument, and a directive application can sit at any of five sites, so the parent is a union of relations and a key names one table. A position is unique in a file whichever relation holds the parent, so those two references are sound and only unspellable, which makes them detections over the union rather than constraints. The keys are why the writers run outermost first and the sweep runs their list backwards.

A value is one relation for four holders, and the only relation here that references itself. An argument of a directive application holds the value passed to it, and an argument, an input field and a directive definition’s argument each hold a default. Those are four sites, which by the rule above would be four relations, and they are one. The rule is about what a key can defend rather than about sites: a directive’s parent is one kind of node at each site, so a key is available and splitting buys one; a value’s holder is four kinds at every site, so no key is available at any grain and the split buys nothing. What a value does have a key for is the value enclosing it, which is always a value, and that reference is a real foreign key back into the same relation. Two consequences worth knowing. The writer inserts a level at a time, shallowest first, because a batched insert of a whole tree gives the constraint no order and a child can be checked before its parent is there. And the holder position is repeated on every node of a tree rather than held at its root, so one predicate fetches a whole expression and no reader climbs to find out which slot it was written at. An object’s field gets no row of its own: a field is a name and a value, so the value is the row and the name is a column on it.

Nothing here references an anchor. The anchors are derived from these rows, so an entry holding a foreign key into one would make the entries conditional on something built out of them.

No node carries an ordinal. Within a parent, ordering the nodes of a kind by the position they were written at is total, two nodes being unable to share a position, and it is the order the author wrote them in. A column would restate the key.

The file is named twice. source_name sits in the primary key carrying no foreign key, and a nullable source_ref beside it carries the reference under CHECK (source_ref IS NULL OR source_ref = source_name) and FOREIGN KEY (source_ref) REFERENCES store_source (source_name) ON DELETE SET NULL. One column cannot be both a key that refuses a delete and a twin that survives one; H2 resolves that ambiguity by letting the cascade win. So the key gives up its foreign key, and removing a file from the registry nulls a column on every row that read it instead of deleting them or being refused. The attribution stays in the key; what the null withdraws is the registry’s vouch, and one predicate, source_ref IS NULL, finds both a source removed under a reading and a write that left the column unset. The graph that owns the rows reaps its own on its next refresh, which is what keeps one process from reaching into another partition on a cadence its owner knows nothing about.

The reader does not use graphql-java’s NodeTraverser, which is the obvious tool. It walks the untyped child list, so an interface named in an implements clause and a union’s member arrive as the same TypeName under a type definition, indistinguishable without asking what kind the parent is. getImplements and getMemberTypes are the answer the API already gives. Each relation is populated by its own method naming where its nodes live, rather than by one shared descent, because relations holding different data change for different reasons.

The verdict on the corpus is one relation, and the anchors do not wait for it

graphql_schema_problem holds what went wrong when the documents were read and made into a schema, as the toolchain this graph is built with stated it: graphql-java, the federation link processor and the configured loading rewrites. One relation and not one per stage, because parsing a file, merging the files into a registry, composing that registry the way the generator does and building a schema from it are not four questions a reader has. The parse judges one file at a time; the merge only ever says whether a name was declared twice, three error classes in all; the composition refuses in graphitron’s own words, a second federation @link or a configured tag the @link does not import; the build runs around thirty checks over references resolving, interface contracts, input against output position, directive locations and arguments, uniqueness inside a declaration, and the schema’s own shape. Which one spoke is a column, so a reader who does care can ask without four relations to union.

The verdict is the build’s toolchain’s rather than any SDL reader’s, and the transcription beside it is not: the transcription is reader-neutral, the verdict is that of the toolchain this graph is built with. A reader-neutral verdict is the one that reported 108 undeclared directives on a federated consumer whose build saw none, and a verdict that disagrees with the build is not neutral about the graph, it is wrong about it.

A file the parser rejected is the reason the row carries a site. It contributes no entries at all, so this relation is the only place that can say the file was read, and a reading that dropped the refusal would leave the file looking like one nobody configured. The later stages point at a declaration where they can and nowhere where they cannot, which is why the site is nullable.

Nothing here is re-derived, and that is the point. Computing any of those checks from the declaration entries would be a second implementation of a validation the build already runs, and the two would disagree the first time the library moved.

Distinct, because the same problem said twice is not two problems. One absent type reached for from five fields raises five errors whose sentences are byte-identical: the message names the absent type and the type that reached for it, never the field. How many fields reached for it is a count over the entries, so keeping one row loses nothing.

The anchors do not depend on any of this. They are derived from the entries, so a corpus that does not build still has them, which is the ordinary state of a schema while somebody is editing it: a developer halfway through a rename should not lose every answer the store can give. Going through the error classes, only the collisions contend for an anchor row at all, and there the rule is already settled by reading order, oldest first. Every other class is a broken relationship between declarations that all exist, which is an anti-join rather than a missing row.

Almost no foreign keys, and why that is the design

A document may name a type nobody defines. It may implement an interface that does not exist, and it will not assemble, which is a refusal a reader wants reported rather than a row the store declines to hold. So a reference at this level is only sound where a well-formed file could not break it, which means the graph the row is partitioned by and the file the declaration was written in. Those two are what the entries carry and they carry nothing else.

A reference between declarations is a different matter and stays available where nesting is lexical: a field is written inside the type site that encloses it, in the same file, so a child may reference its parent’s position. What is never a key here is anything a second document has to supply.

Entry facts are what the author said; anchors are those facts met with reality

An entry relation holds a fact the author communicated through the schema. That is what the word means across this store, and the two families spend it differently. A graphql_ entry is what one document said, and its anchor is what the corpus says once the merge has ranked the documents. A graphitron entry is what one directive application asserted, and its anchor is that assertion met with the catalog and the classpath. Both are the same shape: said, then resolved, with the resolving named and kept apart from the saying.

So this stratum’s work is deciding. A directive’s arguments are a syntax; the facts are what the author meant by writing them, and each fact belongs in a relation that holds facts of that kind and nothing else. Normalized, in the ordinary sense, and decided here rather than left as a shape for every reader to interpret. The anchors then have something to check, which is a consequence of the modelling and not the reason for it: modelling for the check directly is how the walk this replaces came to route an input object’s field down its field path, because that was what one consumer wanted.

The question to ask of every column is whether it is an attribute of the fact the row states or a different fact. A different fact gets its own relation. An attribute stays a column, and may be nullable without apology when the definition lets the author leave it out.

A @reference path element is the worked example. It may name a table, a foreign key and a condition method, in any combination, and those are three independent assertions rather than one thing with three attributes: there is no entity here whose attributes they are, which is exactly why a single row for the element needed nine nullable columns and a legality rule no constraint could state. Three relations, one per assertion, and an element making several writes to several of them at one list position. The position is what says they are one step, so composition is a join and every relation stays total.

One tell is worth carrying away from that example, stated as a tell and not as the rule: where a discriminator column’s value implies which of the row’s columns are null, ask whether the rows have one key shape. Key shape is what decides. graphitron_field_reference_step_hop is the second shipped instance, and it splits where the element’s own relations did not: its four arms are told apart by a via literal, and on two of them the foreign key’s name and orientation are identity while on the other two there is no foreign key to enumerate at all, so the element coordinate with the two table triples is already total. One relation cannot carry two key shapes, which is why that one had no primary key and nothing in it refused a duplicate row; it is two keyed tables under a view now, and the view is the only surface the projected nulls appear on. graphitron_field_chain_link_resolution is the same split in three, and the first whose rule cannot follow it: one chain mixes arms, so its walk crosses all three key shapes, and the stage evaluates the rule once and files each row under its shape rather than running one rule per table. graphitron_condition_method_route and its _defect sibling are the other shipped instance, splitting a verdict out of an absence rather than a key shape out of a discriminator.

@asConnection is the counter-case that keeps the rule honest. A page size and an overridden type name are two optional attributes of one assertion, both left out by most authors and neither changing what the row is. One relation, two nullable columns, no split. So is @mutation, which names an operation and may name a table: a mutation with a table and one without are the same kind of thing.

Absence is itself sometimes the fact. A bare @table asks for the table name to be deduced from the type, and a bare @nodeId asks for the target type to be deduced from context; each is something the author said, and it is recorded as the absence of the row that would have named the thing explicitly. That is only sound because the applied-directive row says the directive was applied, so "applied and named nothing" and "not applied" stay distinguishable.

A required argument the author omitted is not an assertion at all. It is a malformed application, it writes no entry row, and what they wrote instead stands in the transcription of the value. A row of nulls would say a third thing that is neither.

Where two columns are alternative spellings of one assertion rather than two, the constraint says so and the relation stays whole: a DATABASE error handler discriminates on a vendor code or on a SQL state and the directive rejects both together, which is one relation and one CHECK.

The names do not say this yet, and that is worth knowing before reading them. The graphitron entry relations carry an ast infix they should not. It is the right word only where the rows really are the parse tree, as the AST entries above are; these are decided facts about what an author meant. The end state is the shape the graphql_ family already has, an entry named graphitron_<x>_entry and its anchor named graphitron_<x>, which is blocked only by the incumbent still holding the unprefixed names.

Enforced by: partly, and the gap is worth stating. A CHECK holds the alternative-spelling case and NOT NULL holds the required one, both in the DDL. SupertypeSignatureGateTest reports when two relations carry the same payload, which catches a split made for a distinction that turned out not to be one, and it reads as a census of which facts are the same fact stated at different sites. What nothing checks is the rule itself: no gate can see that a nullable column is standing in for a fact, because the evidence is in what the readers do with it rather than in the schema. The symptom to look for is a reader discriminating on IS NULL over an entry relation; where one appears, a fact was left undecided at capture and every reader has been deciding it since.

Ownership: the second axis, and why intent_ dissolves

Stratum answers what a row is a function of. It does not answer who is responsible for keeping the row current, and that is a different question with different consequences. A relation has both answers, and intent_ today conflates them: the family is named for the first and is being used to carry the second, which is why it holds rules that have no business being derived at read time and rules that could not be anything else, under one prefix that cannot tell them apart.

The rule for the second axis: a view that reads relations from only one family is a view of that family. It is filed there, named there, and owned by that family’s gatherer. graphql_ and graphitron_ are two families and two gatherers, but not one each, and that is the rule working rather than an exception to it. A family names a vocabulary; an owner names a writer, and meta_relation has carried the two as separate columns from the start. The SDL walk owns graphql_ whole and owns the as-written half of graphitron_ beside it, decoding each application into graphitron’s vocabulary while it stands on the coordinate: those rows are a function of one document, which is what makes the walk their owner under this very rule. The graphitron gatherer owns the resolved half, the stages that join an entry against the catalog or against each other, and it runs after every crawler has flushed because that is the earliest moment its inputs exist. Two vocabularies over one document, then, rather than two passes over one corpus. What lets a graphitron_ view read the transcription it derives from is not that the two are one family but that the roster grants it: meta_gatherer_dependency gives the graphitron gatherer sdl and catalog, and MetaDeclarationGateTest.aDeclaredViewReadsOnlyWhatItsOwnerMay holds every declared view to its owner’s set. A prose claim about shared ownership is what that machinery replaced. store_graph_source is the spine every graph-partitioned relation joins and counts as neutral.

The rule is about filing, not about form, and the distinction is the whole of it. A gatherer reads its corpus as a stream and writes what it has in hand; a view sees the corpus whole and can ask questions no streaming pass can. Both are real powers and the rule gives up neither. It does not say that a single-family rule should be flattened into the walk that feeds it. It says that rule belongs to the walk’s family, whatever shape it then takes.

What ownership buys is refresh authority, and that is where the leverage is. A gatherer knows when its own family is complete, because completing it is what the gatherer does. Each gatherer sweeps its own rows from a list it holds itself, so a relation moved between owners moves its lifecycle with it, and the wholesale clear that once stood outside them is retired. Which is what lets the configuration gatherers run before the pass instead of after, and so lets a build the pass refuses still say why. The declaration is a build-time gate on that arrangement rather than a mechanism inside it, and the distinction is worth stating because an earlier revision of this page had the clear consulting meta_relation to decide what to leave alone. It does not, and nothing else that runs reads that relation either; MetaDeclarationGateTest is its only reader. What the declaration buys is that a new relation cannot arrive without an owner, which is a real thing to buy and not the thing the sentence it replaces claimed. A relation whose refresh is known to somebody may be stored by a stage, denormalized, indexed, or left a view entirely, at its owner’s discretion and without anyone else having to be told. A rule that moves into a family therefore does not get converted into anything; its owner decides its form.

So intent_ is not "the derived family", which describes how rows arose and separates nothing. It is the shape a pipeline takes when a fact is not written down, and it goes away entirely. Every relation has an owner and there is no unowned relation: a rule reading one family is owned by that family’s gatherer. A rule whose facts appear to cross families is a rule with a single-family part inside it that has not been separated out, and the cut is decidable rather than editorial: expand it through every intent_ relation it names until only captured relations are left, reading a stage-written table as the rule its stage states so that a stored table cannot hide a crossing underneath it, and ask which family each captured relation sits in. Every part lands in one. A part that seems not to is a fact nobody wrote at a grain, and the answer is to write it at that grain in the family whose corpus it comes from, as early as it can be written, where every reader reaches it and no reader derives it again.

An earlier reading of this section made the crossing rules permanent and gave them an owner of their own, a derivation gatherer running after every corpus gatherer. That reading is retired. It conceded the premise that some rules must be composed late, and late composition is the thing this page is arguing against: the graphitron gatherer already runs last, after both crawlers have flushed, as stages in an order it chooses with the whole transcription and the whole catalog in hand. A rule it can compute in a stage needs to be neither a view nor a reader’s join. The relations this arc has reached have been deleted rather than refiled, which is the evidence that the residue the earlier reading reserved a gatherer for is not there.

That is also what settled the materialization register. The store once carried a register scheduling refreshes for rules nobody owned, a mechanism standing outside the ownership rule and answerable to nobody, and its roster showed exactly that: every registration was an intent_ relation, and not one sat in a family with an owner. It was not a scheduling facility that happened to be used here. It was the lever a reader reached for when the fact could not be moved earlier, and the fact could not be moved earlier precisely when nobody owned where it would go. So the question was never "should this registration exist" but "which family should have written this down, and how early". Answered for every registration, the register had no subject and was dropped: each of its rules is now a stage in the graphitron gatherer’s derivation stratum, writing a graphitron_ table from the rule it keeps beside it as a stored _rule view, and DerivationStratum lists those stages in the one order they may run.

Pushing a part down does not convert it into anything. It usually leaves the view form behind: a stage of the owning gatherer computes it once, with its inputs in hand, and writes it. What a reader then names is a relation in the family whose corpus the fact came from, at the grain the reader wanted, which is what "available to as many readers as possible" means operationally. Nothing a reader names disappears in the move; what disappears is the prefix and the second derivation of the same thing.

The shape exists in the store already, unevenly applied, which is what makes this a correction rather than a proposal. graphql_element is the supertype-over-subtype-relations shape done correctly, and it is a table rather than a view because that is what the shape requires: it carries the coordinate as a primary key, its element_kind says which subtype a row is, and the six graphql_*element relations key into it while each keeps its own grain. A view could not do that. A supertype exists so its subtypes can reference it, and a foreign key cannot name a view, so a union standing in for one buys the vocabulary and none of the integrity. ClassificationDomainCapture is a gatherer-owned derivation, its own javadoc calling it the SDL gatherer’s rooted traversal at its last stage, writing inside the capture transaction; its output is filed under intent, which is precisely the misfiling this rule names.

What the store does not have yet is per-gatherer transaction control, and it is the one genuine prerequisite. The capture pass runs every gatherer inside a single transaction, so no gatherer can commit its own family and then refresh its own relations against statistics that reflect what it just wrote. H2 commits the current transaction as a side effect of ANALYZE, which is the mechanical reason. One exception is already carved out for exactly this: a capture into a store no derivation-stratum table holds a row in commits its facts and runs the stratum outside that transaction, each stage committed and analysed before the next one plans, so every stage can be planned against statistics it has. What this rule asks for is that exception promoted to the normal shape, one boundary per gatherer.

Enforced by: the declaration half of the machinery, progressively, and by the family saying so itself. Every intent_ relation whose comment is not its declaration rendered opens with a notice that the family has no owning gatherer and is being retired, which jOOQ copies into the generated javadoc and the schema reference renders, so a contributor meets it at the relation rather than only here. meta_relation binds each declared relation to its owner, meta_gatherer_dependency declares whose rows a gatherer may read, and MetaDeclarationGateTest.aDeclaredViewReadsOnlyWhatItsOwnerMay walks each declared view’s stored definition with ViewReferences and denies a read outside its owner’s declared dependency set. The gate binds per declared relation, so its reach grows with the declaration migration and covers nothing that migration has not reached. What stays unenforced is the misfiling census itself: nothing yet fails when a rule reading one family is filed under intent_, because that check is a query over the declared owners of the intent_ residents and those declarations land last, in the roster order the declaration item states. Stating the count here instead would be the unguarded census this page warns about two sections up; the roadmap item carrying this work holds the current figures and this page holds the rule.

Derived reads are views, not stored facts

Most of what a consumer reads is derived, and each recurring derivation is easy to mistake for a stored thing and model wrong. The candidate space (completion, enumeration) is a base relation read without its key constraint, the same relation a resolve reads with one; storing a candidates list would duplicate the relation it selects from. Reverse indexes (find-usages, what-is-at-this-cursor) are inverted functional dependencies maintained for read speed; that a reverse lookup can find an entity does not make the lookup’s input the entity’s key. Stable ids are rendered keys, per the key discipline above.

The converse test matters as much: two spellings of one value are two base columns when neither is a function of the other, and the way to settle that is to look for the case where the derivation fails rather than the cases where it works. A method’s return type is carried both erased, as the JVM descriptor spells it, and as the source declared it. Most of the time the declared form determines the erasure, which makes the erasure look like a view over it, and one example refutes that: a type variable declares T and erases to its bound, so the declared form does not name what the erasure is. The other direction fails just as plainly, List not naming the element type List<Film> declares. Both columns are base, and the DDL comment owns that argument at each of them, because the next reader will otherwise reasonably try to delete one.

Which form a reader takes is then a property of the question, not a preference. A surface spelling a signature for an author wants the declared form, because the author is reading their own source back. A check on whether a method returns a particular type wants the erasure, because every Field<X> answers that question the same way while the declared spellings all differ. Storing one and computing the other would have forced one of those two to be wrong.

What a single column may hold has a rule of the same kind, and it is two clauses rather than a distaste for renders. A column may carry an opaque value the store did not compose, a captured message or a transcribed docstring, provided nothing joins, groups or filters on it; the DDL grants that permission in the same words at each column that has it, "display material, never a dimension". A column anything does join, group or filter on must be atomic to the engine and a function of its relation’s own key, and what fails that clause is a collection inside a scalar, with or without a delimiter to make it obvious. intent_type_backing_conflict holds both verdicts under one key: candidates, a count over the classes contesting a type, passes, and the same set joined into one string would fail, its element grain sitting inside the value where no key, constraint or join reaches it and only string surgery gets it back. What the rows answer and the joined string cannot is membership. A serialized set answers equality of the whole set and nothing else, where the contesting classes as rows on intent_type_backing under that same key answer that and membership too, for one join, which is what makes a set used as a canonical group key the case that looks legitimate and is not. That difference was paid once rather than being hypothetical: the claim-conflict relation used to carry the claiming directives of a conflict as one serialized value, the MCP diagnostics surface offered it as a filterable dimension, and asking for the conflicts involving one directive answered with only the conflicts whose entire set was that directive; the surface joins and asks membership now. Order is the neighbouring case and splits the same way, admissible as data (a captured ordinal, a position column) and inadmissible as a rule copied from a consumer’s vocabulary with nothing binding the copy, which is the argument the provenance section already makes about a naming order that is no captured fact of any graph and so no view’s to express.

Enforced by: CollectionValuedColumnGateTest in graphitron-model, for the named aggregate constructs in the DDL’s statement regions, reading the authored text rather than a booted catalog because a booted store has already lost the spelling. The gap is disclosed in that gate’s own javadoc rather than left to be found: a row-local scalar expression that discards part of a value trips nothing, detecting serialization inside an arbitrary expression not being mechanizable, so this paragraph is the coverage for that residue. It is closable case by case, each such expression being findable in review, and its exemplar is the path-truncation column the diagnostics surface used to carry.

A derivation gets a relation as soon as a second reader asks it. Written inside the first reader’s query it is a CTE, which is fine while it has one reader and wrong the moment it has two: the alternative to naming it is the second reader re-spelling it, and two spellings of one resolution agree exactly until one of them changes. intent_bound_table is the shipped case. Which catalog table a type’s @table binds to began as a CTE inside intent_column_match_claim, the classifier that asks it on the way to a claim, and became a view of its own when the language server started asking the same question with no claim in view. Naming it also made its declined case answerable: the classifier wants exactly one candidate and refuses an ambiguous binding, an editor wants to offer every candidate, and both readings come off the same rows because the arity is a column rather than a rule inside whoever counted first.

Resolutions layer among themselves, and the layering axis is what a resolution is keyed on. A binding is keyed on a coordinate, a type or a field, because that is what a reader holds when it asks. Underneath sits the rule that does not vary by coordinate at all: graphitron_spelled_table answers how a written table name meets the catalog census, and a qualifier bound against the schema, an unqualified name matching case-insensitively and a membership scope over the graph’s catalog sources are the same rule whether the name was written in @table(name:), in a @reference path element, or as a `@mutation’s delete target. The qualifier arrives already split: capture partitions a written reference on its first period and stores both halves beside it, because a grammar is a parse boundary SQL cannot express, so the rule reads a partition rather than performing one. So the binding view became a keying over the spelling relation rather than a second copy of it, and the population of that lower relation is every spelling the graph authors, which is the honest reading of a relation keyed on a string: not every name the census holds, and not the subset one site happens to write.

A case fold is a stored column, and it is minted only where an authored spelling meets a catalog name. That crossing is the whole reason one exists: an author types a GraphQL or SDL identifier, the catalog holds a SQL one, and the two namespaces disagree about case, so each side of the comparison carries an _upper companion the database computes and nobody writes. Two consequences follow and both are load-bearing. A comparison between two values of one family mints nothing, because there is no boundary to bridge; where such a comparison does want folded operands it joins the relation that owns the spelling on its key and reads the fold there, which sql_name_matched_key_column and intent_field_reference_discovery each do. And a derived view never forwards a fold to its readers, for the same reason it does not forward anything else it merely carried: the fold is a property of the relation that owns the spelling, and a reader wanting it re-joins rather than having a column threaded through every view between. What that buys is one rule with no list attached, and views that hold no per-row case fold on any comparison the rule reaches. What survives it is the one comparison the rule declines to serve: the defect view over node-identity metadata matches a generated table class’s stated key-column name against the catalog’s own column names, and both of those are values the crawler produced, so the fold there is a hedge rather than a semantic. That is not a fold this rule owes a column to; it is a fold nothing owes anything to, and it goes away by becoming exact rather than by being stored.

The same discipline splits a derivation when only part of it needs recursion. A @reference path resolves sequentially, an element departing from where the previous one arrived, so the chain is a recursive walk; but each element’s own resolution has no recursion in it. Those are two views. graphitron_field_reference_step_hop enumerates every table-to-table hop an element could express, both orientations of its foreign key included, because which direction the element means depends on where the chain stands; graphitron_field_reference_step_target walks them from the enclosing type’s binding. Keeping them apart is what lets the recursive term be a single join instead of a copy of every element arm, and it puts each arm’s rule in one place. The hop is itself two relations under that name, for the reason the column question above gives rather than for this one: its arms have two key shapes, so graphitron_field_reference_step_hop_keyed holds the hops a foreign key identifies and graphitron_field_reference_step_hop_keyless the hops none does, each keyed and refilled from its own rule, with the name every reader spells a view unioning them. Two splits of one subtree on two different axes, and neither implies the other: recursion is why the walk is not the hop, and key shape is why the hop is not one relation. Splitting the chain also forced two arities apart that a single count would have conflated: a path element naming a table three foreign keys connect reaches one destination by three routes, so targets and candidates are separate columns, and a reader needing only the table can trust an answer a reader rendering the join cannot.

A recursive view has to terminate on the population the store can hold, not on the one the subject would have. The worked example is a view that no longer exists, retired for what measuring it showed: an all-pairs assignability closure over the census’s declared supertype edges, taken as a view because a Java class hierarchy is acyclic where the SDL type graph intent_type_domain closes over is not, and materializing is what a closure over a cyclic relation costs. Acyclicity of the subject is not what makes such a recursion safe. The relation the closure ran over is the whole classpath census, which spans every entry every graph ever read, so one class name declared by two entries is routine and two such names declared into each other are a cycle both classfiles can be valid under. Three measured facts came out of that, and each is a rule for the next recursive view rather than a fact about the retired one.

A termination guard is unavoidable, because H2’s recursive UNION does not deduplicate against rows earlier iterations produced. A cycle under it is not a wrong answer but a hang with no diagnostic. The schema’s two reference-chain closures escape the guard a different way, on a strictly increasing position that is a column of the data rather than a predicate over the recursion’s own history.

A path guard is not free, and this is the claim the retired view got backwards. Guarding on the path terminates and then enumerates simple paths, which over 8,821 declared edges across 153 classpath entries costs seventeen seconds. That census declares no class name twice, so nothing about the seventeen seconds is a cycle or a duplicate; it is what an all-pairs closure over a census-scale relation costs. A closure that has to be cheap is anchored at the names its consumer asks about, which over the same census answers in a second, and it grows with what the consumer asks rather than with the census.

And the hazard the guard was reached for is duplicate rows more than cycles. A recursive term that joins on one column and projects another turns two rows differing only in the projected column into identical output rows; UNION ALL keeps both, both recurse, and the frontier doubles per hop. Forty stated rows reproduce the hang, so this needs no census-scale fixture to show. The rule that follows is to recurse over the pairs the edges denote rather than over the rows that declare them, which is what intent_authored_field_claim’s lookup-bearing closure does by recursing over an `input_object_field_edge CTE instead of over graphql_field directly.

The general form: the argument for a recursive view is about the rows the relation can contain, and a subject-level invariant ("a class hierarchy is a DAG") is only the same argument when the relation holds exactly one subject.

A derived view carrying a window function or a recursive term cannot be pruned by a predicate applied outside it, so a reader takes it once per answer and pairs it on its key rather than correlating it per row. The measurement that settles it, from the MCP schema read: intent_column_match_claim and intent_resolved_field_claim read as correlated MULTISET subqueries under a per-field projection cost twenty-four seconds over sixty types and eight hundred coordinates, where every other field-grain slot together cost a third of a second and those same two views read whole cost between two and seventy-five milliseconds. The column-match classifier collapses its matches with a ROW_NUMBER() OVER (PARTITION BY …​) over a derived relation, and a window sees its whole partition whatever the outer correlation says, so the correlated form pays the entire view’s evaluation once per driving row. Driving the statement from that view instead, filtered to the page, with the witnesses joined in as arity-preserving left joins, brought it under two seconds. A base relation correlated per row is the opposite case, an index seek that nests freely. So the rule is narrower than "avoid correlated subqueries" and sharper than "measure it": the view’s own shape decides, and a derivation this deep wants to be first in the FROM clause. Each view owes that warning in its comment, because the cost is invisible at the call site.

H2 inlines a view wherever it is named and eliminates no common subexpression, so a relation a derivation names four times is evaluated more than four times, and the multiplicities compound down a tree of views. Nothing at a call site shows it: a reader sees one SELECT against one name. The static form of the count is a reported metric, report-inline-multiplicity in the roadmap-tool run, which counts each relation’s textual references in each view body and multiplies them down the tree, needing no database and no profiler. What it ranks is breadth, and breadth is not cost: 2528 namings of a relation that answers in 0.4 s is not a problem, and 83 namings of one that answered in 20 s was. So the metric names suspects and a per-relation timing prices them, which is why it reports rather than gates.

An anti-join against such a view is the case rewriting cannot fix. An anti-join has no drivable side: the excluded relation is probed per candidate however the statement is written, and five spellings of one exclusion, from a correlated NOT EXISTS to a RIGHT JOIN driving from the view, measured within a few percent of each other. Materialization is the only lever, and it applies to the relation being expanded rather than to the expensive relation inside it. Measured on a condition-site view since retired, which excluded what a companion view claimed about which parameter positions receive the source table: snapshotting the recursive closure underneath left the read an order of magnitude above the floor, because the view around it was still expanded per row; snapshotting the excluded relation itself hit the floor. That reader later took a third lever neither of those is, and it is the one to reach for first where it is available: the excluded fact was a question about a Java method, so the gatherer that reads the classfile answers it and stores the answer as a column. The anti-join is a column predicate now and the companion view is gone. Materialization is the lever when the excluded relation states a rule over rows the store already holds; capturing the decision is the lever when the rule is really a question about a source nobody has asked yet.

A derived relation joined on an expression rather than on a column is evaluated once per driving row instead of once. Measured against the sakila example’s schema and catalog, a resolution rung shaped exactly that way over graphitron_resolved_type_binding cost 19.9 seconds for 157 rows, where projecting the expression as a column in an inner derived table first and joining the binding on that column cost 0.13 seconds for the same rows. Three controls on the same fixture isolate the expression as the term. Joining the field’s own named-type column directly, wrappers ignored, is 0.08 s, so the cost is not the row count. Routing the expression through graphql_type and joining the binding on its column, which is how two other sites in this schema spell it, is 19.6 s and no fix at all. And materialising the inner relation in a WITH clause first is still 19.6 s, because H2 inlines a non-recursive WITH exactly as it inlines a view.

That last control is the one to read carefully, because two rules stand next to each other here and read as contradicting each other when they do not. Extracting a relation for tidiness changes no join key, so the expression reappears in the key of whatever named the extraction and nothing is bought. Extraction that projects the expression as a column and joins on that column is the two-orders-of-magnitude fix. `graphitron_argument_scope_table’s own comment is the live exemplar of the second, calling its inner derived table load-bearing rather than a formatting choice, so it is the wrong citation for the first.

Every figure in this section is a timing, and one further engine behaviour decides when such a figure is real at all. H2 reuses the result of a repeated identical query, so a repeat is not a repeat until OPTIMIZE_REUSE_RESULTS is off. Measured on 2.4.240, six runs of one view query cost 98 ms and then 0, 0, 0, 0, 0, and the per-statement statistics H2 keeps when asked reported an average of 15.4 ms for a query that takes about 92. The under-report is proportional to how many repeats were asked for, which makes it worst exactly where an author is being careful and averaging several runs to damp noise, and it is silent, a plausible small number rather than a missing one. The setting is database-wide rather than per connection, so one statement covers every reader a session mints, and it has to be set the same way on both sides of any before-and-after comparison, because disabling reuse moves the absolute figures and not only the repeats.

A second engine behaviour decides when a plan is real, and it is the one that catches an author who has learned the first. EXPLAIN without ANALYZE is attractive precisely because it does not execute, so it reaches statements no timing can, and on this schema it does not render the plan that runs. H2 optimizes a view’s inner query per set of index conditions, at execution, and a bare EXPLAIN renders the unmasked form. Measured on 2.4.240 over the twenty refresh statements the materialization register then carried, comparing a store with statistics against the same store with none: plain EXPLAIN reports three of the seven statements whose plan actually moves and agrees on the other four, whose rows visited move by factors of three to eight. EXPLAIN ANALYZE renders what ran. Two consequences for a before-and-after on plans. Compare the executed plan, and compare its text rather than a reduction of it: a reduction to the relation order and the index names, which is the obvious one to write, missed four of the same seven, because the shape that changes is one index used two ways and the difference lives in the seek conditions. scanCount is the only rendering in these plans that varies with statistics by construction, so removing that annotation and comparing the rest is both simpler than a reduction and strictly more faithful. The test that held that comparison left with the register, and the rule stands for the next comparison anybody writes.

A recursive term re-evaluates a relation named in its step once per accumulated row. The recursive UNION joins its own accumulated output against the step’s input, so a relation named there is evaluated as many times as the walk has rows rather than once, and inlining it makes the whole subtree under it the thing being re-evaluated. Measured on the same fixture: a step naming graphitron_node_id_decode_hop_column, which is 6.8 s evaluated on its own, produced 20 rows in 146 seconds, which is that cost times the rows the walk accumulates, where the same walk over those rows as a table is a little over 3 s for the whole reader. What was not the term is the step’s join predicate: collapsing its six-column coordinate key onto the use site, on the theory that four null-safe disjunctions were what stopped the step being planned, is a real simplification that moved the timeout not at all.

The general form tying those together, and the sentence to reach for before writing a rewrite: what makes a relation expensive is being a view that something reads many times, not how the reader spells the read. Two predicate rewrites over the node-identity family bought exactly nothing for that reason, and the two changes that did pay changed what a relation is rather than how it is named, one turning a join key from an expression into a column and one turning a recursive step’s input from a view into a table.

A reader meeting that rule reaches for CREATE MATERIALIZED VIEW, and on H2 it is unavailable rather than merely unattractive. H2 2.4.240 has the statement and real snapshot semantics, and on a synthetic stack of the shape above it takes a per-graph sweep from 5.9 seconds to 4 milliseconds, so the reach is well motivated. Four defects stand in the way, each a defect rather than a preference, and the first of them is also what makes this prohibition need no gate of its own. While a materialized view exists anywhere in the database, reading INFORMATION_SCHEMA.COLUMNS throws an internal NullPointerException, because the view is a shell object whose inherited column array is never populated and every query against its name is rewritten to a backing table before anything else notices; StoreCatalog reads that relation and jOOQ codegen boots off this schema’s live metadata, so one materialized view anywhere breaks the model build and the generated schema reference. A persistent database containing one cannot be reopened at all, because the stored DDL replays as CREATE FORCE MATERIALIZED VIEW and H2’s parser answers FORCE on that statement with an unimplemented-operation throw; the fact store is file-backed and persists warm across builds, so this one rules the feature out on its own. Dropping one leaves its backing table behind and reports a BASE TABLE row with a null name, so re-creating the same view then fails. And refresh is manual with no way to derive an order for it: refresh does not cascade, so an outer view stays stale until refreshed in its own right, and H2 exposes no dependency information anywhere, the three SQL-standard view-usage catalogs being absent and jOOQ’s own metadata carrying no dependency accessor. A refresh chain over layered views would therefore be hand-maintained ordering, which is the shape SchemaIdentifierDriftCheck exists to refuse: the universe of relations comes from the booted store and never from reading the DDL. The derivation stratum’s step order does not contradict that objection, because the objection was to an ordering with no derivable source: the stratum’s list is kept by hand, and StageOrderGateTest parses each stage’s rule out of the booted store’s own stored view definitions and fails the build on a step placed ahead of a table it reads, so the list is checked against the catalog rather than trusted.

The ruling is conditional on that being the state of the tool rather than on the feature being wrong. Each of those defects was traced to its cause and the blocking ones were fixed and verified against H2 trunk, in twelve lines for the metadata read and about sixty for the persistence, at which point a file-backed store survives restart with its snapshot intact and correctly stale. None of that is in any release, none of it is filed upstream, and 2.4.240 is the latest, so there is no version to move to. Revisit the ruling when a release carries the fixes, and not before. Note also what materializing would and would not buy even then, because it is easy to over-read: batching a read pays one evaluation and materializing pays one evaluation plus the refresh, so for a write-then-read-once workload they are equivalent, and a snapshot only pays off where a relation is read many times between writes.

So where a derivation genuinely must be paid once and stored, the reduction is an ordinary table populated INSERT INTO derived SELECT …​ FROM <view>, which keeps the view as the single statement of the rule while making reads a plain indexed scan, and is the shape every stage in the derivation stratum takes, and leaves a normal table for every other purpose: indexable by its own name, cleanly droppable, visible to codegen on our terms. Per-reader, the same shape is a LOCAL TEMPORARY table that disappears with its connection. Say LOCAL explicitly and treat a bare CREATE TEMPORARY TABLE as a trap: it defaults to GLOBAL, and H2’s global temporary tables share their rows across every attached session rather than only their definition, so in a store the language server, the MCP server and concurrent module builds all hold sessions on, one reader’s scratch snapshot would be every reader’s.

Which lever to reach for is itself ordered, and the top of that order is a modelling argument rather than a cost one. A captured fact is the top rung, and the reason is not that it is cheapest: a rule that reconstructs what capture could have written is a defect in the model whether or not any reader is currently slow. A set of subtype relations with no relation for the thing itself is that defect, and the reader cost it produces is a symptom whose absence proves nothing. Argued from cost alone this rung cannot tell an author to write the supertype for a relation nothing is waiting on, which is exactly the case where writing it is cheapest and the only case where it can be done without moving a reader. That it is also the cheapest rung is a consequence and not the argument: where the value is something a walk already read, writing it down at capture leaves every reader joining a column and no derivation either to evaluate or to refresh.

An index on the join key is the second rung, and a rung of its own because storing a rule has been delivering it silently: a stage-written table is a table, a table is what an index can sit on, and the measured gain reported below on the reference-step hop table is half index and half statistics rather than storage, so a figure this page attributes to storing a rule may be an index nobody priced separately. A key counts as the index for this rung where it leads with the columns the reader joins on, which is what that relation carries now: it is two keyed tables under a view, both keys leading with the eight columns the walk seeks, and the declared index the figure below was taken with was dropped on a measurement that found the keys serving the seek for nothing. And where the key is an expression rather than a column it is not reachable at all until the expression is stored as one, which is the join-key rule above read as a lever rather than as a hazard.

A rewrite is the third rung. It usually changes nothing the planner cares about, which is the general form above read from the other side, and the exceptions are the shapes this page names rather than a licence to try: a correlated probe into a body carrying a window, which no outer predicate can prune however narrow the question, and a join on an expression no index can serve.

Storing a rule is the last rung, and it is last because it is the only one of the four that adds work rather than removing it. What it claims is that the rule is right as a view and only too slow to evaluate per naming, which is why the rule stays a stored _rule view beside the table and the stage is one INSERT over it, where a hand-written producer argues in its own table comment that no view could express its rule at all. The trade it has to win is the one stated with the prohibition, a write per capture against the re-evaluations it avoids, so a relation read once between writes gains nothing from being stored and a relation read many times gains the difference. It also carries a precondition the three rungs above it do not, because it is the only one bought on behalf of readers rather than of the model: a stored table no consumer reaches is a write paid on every capture for nobody, so whether the reader exists is a question to settle before reaching for this rung and not after. And it is filed like any other relation: the stage belongs to the gatherer that owns the family, so storing a rule never needs a mechanism standing outside the ownership rule.

Storing a rule also does something the three rungs above it are usually reached for instead of, and knowing what it is changes when it is the right lever. H2 inlines a view at every naming and performs no common-subexpression elimination, so the statement the planner has to build before it reads a row grows with the depth and the fan-out of the view graph above it, multiplicatively. That size is a function of the DDL alone and can be counted from it: give a base table one, and give a rule one plus, for every relation name it spells, that relation’s count times the number of times it spells it. A stage-written table counts one, which truncates the tree at that name for every rule above it. Counted on this schema when the figure was first taken, the largest statement any relation reached was 963. Counted with every stored rule demoted to its view, the largest was 2739455, and two relations at the top of that range exhaust a four-gigabyte heap while still parsing, with no execution frame on the stack at all. So past some size storing a rule is not buying speed, it is buying a plan existing, and that is a different argument from the write-against-re-evaluations trade above. It is also an argument the top rung answers better: a fact capture writes is a table too, and truncates the same tree with no derivation to pay for.

Having reached for that lever, the relation to store is the one every expensive reader has in common, low enough in the derivation tree that storing it stops the re-evaluation for all of them, and not the relation that looked slow from where the reader happened to stand. intent_node_id_decode_endpoint is the measured case, taken while the register materialized it: three relations read it and each was paying for its whole subtree, so one 5.4 s refresh of it took the two that were timed from 7.5 s to 2.4 s each. It is a view again today, so these figures record what the lever bought when it was pulled rather than the store’s current shape. The counter-case is the same test read the other way, a candidate whose one reader nothing exercises yet having every refresh buy nothing, which is what the register’s hop-column row said about itself and why it was made in the increment that added the reader rather than when the cost was first seen. So the test is to count the candidate’s readers, price its refresh, and prefer the deepest relation whose storage removes re-evaluation for more readers than the one you started from. Stopping short of that depth is measurably not a fix: with the endpoint stored and the recursive step’s input left a view, the recursion reads a 2.4 s view once per accumulated row and is no better off than before.

What the depth rule leaves out is that storing a rule is a shared investment, so the reader it was bought for is not the only reader whose cost it changes. Storing a relation replaces a rule with a table, and a table with no key on it is a heap: every join a derivation performs against that target scans all of it, where the same join against the rule reached base relations the planner could seek into. So the cost lands inside the derivations that read the target, once per driving row where the reader correlates and once per iteration where it recurses, and not on the reader’s own predicate. That distinction is the one to hold onto, because the reader’s predicate is where it looks like it should land and is not: a reader whose derivation carries a window function or a recursive term was never pruning the rule from outside in the first place, by the rule stated earlier on this page, and such a reader still gets an order of magnitude cheaper from an index on the target. Nothing about the motivating measurement would show any of this, the cost being invisible at the call site.

The lever is therefore underneath the reader, which is the ordinary direction, and no reader has to be restructured to reach it: declare an index on the target, on the columns a named reader joins it on, and give the planner current statistics after the stage that fills it. Measured on the scaled fixture the read-cost gate of the time ran on, that took the deepest reader of the reference-step hop table from 18308 scans to 523, and it removed all three of the large regressions that stood in that gate’s pinned set. The example is kept at those figures and the declared index it was taken on is gone, which is the rung’s own point read once more: that relation is two keyed tables now, their keys lead with the same eight columns, and dropping the index moved neither the scan counts nor the plans, so what the rung buys is a seekable ordering on the join key and not a CREATE INDEX in particular. Half of the gain is the index and half is the statistics, and the second half has a placement constraint worth knowing before reaching for it: H2 commits the current transaction as a side effect of ANALYZE, so it cannot run inside a transaction a capture must not commit. Which transaction that is depends on the store. A capture into a store whose stratum tables already hold rows runs the stratum inside one transaction and analyses after that closes; a capture into a store no stratum table holds a row in runs it outside, one committed transaction per stage, analysing each table a stage wrote before the next stage plans. The paragraph below is why the second one exists.

That placement has a consequence for the stratum itself, and it is measured rather than inferred. Worth knowing before reaching for a plan comparison on a cold store, and worth restructuring one capture for, which is what the cadence at the end of this paragraph is. Every statement a single-transaction stratum issues on a cold store is planned with no selectivity on anything it reads, while every timing anyone takes of those same statements is taken afterwards, against a store the analysis has run on. Those are not the same plan on a subset of the stages, and the mechanism is the paragraph above arriving where its own fix cannot reach: without statistics H2 assumes each column has half as many distinct values as its table has rows, which reads a partition column as highly selective, so a one-column seek on graph_name prices as though it were nearly exact and wins against the multi-column index the settled store picks. One statement of that column the store does not wait for a measurement to make: a partition column’s distinctness is a fact about the shape of the model rather than about a population, so every base table carrying graph_name is declared SELECTIVITY 1 when the schema is created, which no.sikt.graphitron.model.catalog.GraphPartition holds and the DDL’s own header states. It is the only statement of that column a pass planning inside a transaction can have, ANALYZE being a commit. What it is worth depends on what else the target carries, and the two levers are worth reading together: on the single reference-step hop table the declaration alone took the decode-hop rule’s read from 10939 rows visited to 1113, and on the keyed arm tables that replaced it the same read is 1499 against 1105, the key having taken most of the same cliff. The cold regime below is therefore a store carrying that declaration and nothing analysed, which is what a run meets. What the plans turned on, measured while these were register refreshes, was the stored tables' statistics rather than the base fact tables', which is the half no single transaction can supply, the tables being what the pass itself writes: analysing the fact tables alone reached none of them, and analysing the stored tables alone reproduced the settled store’s plans on every refresh. The plan-text test that held those three claims left with the register, so they stand here as a measurement. The gap this opens grows with the schema rather than staying a constant factor, a per-driving-row seek being linear in driving rows: on the read-cost gate’s fixture scaled to 187, 349, 673, 1321 and 2617 fields, one whole refresh cold against the same refresh with the targets analysed ran 338 ms against 274, 504 against 377, 1841 against 976, 9633 against 3036, and 68.3 s against 13.2. Read the analysed column too, because it was a separate finding about the register rather than about statistics: thirteen seconds for one refresh of a 2617-field schema, from 274 ms at 187, is superlinear before any statistics question is asked. And it arrives larger than the fixture says at consumer scale: on a captured store from a real consumer schema, the measured prefix of one cold refresh cost 6293 s against 90.8 s with the targets analysed, sixty-nine times, and the whole refresh of that capture took four hours and nineteen minutes while an ANALYZE over its targets cost 0.2 s. What the tree does about it is a cadence rather than a rewrite. A capture into a store no stratum table holds a row in runs the derivation stratum outside its load transaction, one committed transaction per stage, analysing what the stage just wrote, so every stage plans against the tables the stages before it filled. That store is both the case that hurts, every stratum table on it being empty, and the case where committing between two stages empties nothing that was committed; every other capture keeps its single transaction. The condition is asked of the stratum’s own tables and not of the store’s graphs, which is a correction rather than a wording: it read store_graph as a proxy at first, and three writers mint that anchor row for reasons of their own, the build’s own run-configuration capture among them, so every consumer build found a row on a store whose every target was empty and took the in-transaction cadence with no statistics anywhere. DerivationStratum.analysingCadenceApplies holds the predicate and the class names the two conditions the cadence rests on. StagePrerequisiteStatisticsTest holds the invariant in both directions: a cold capture, a warm one and the analysing cadence run directly each meet every stage’s prerequisites analysed, and the one-transaction cadence on a store with no statistics meets every one of them unanalysed, which is the control that shows the instrument can see the difference. It asserts statistics rather than a wall clock for the reason this page gives about figures in a test tier.

So the cost of storing a rule is not only what it saves its own reader but what it does to the others. While the register stood, a build gate priced every pair of a registration and a relation reaching its target, in both shapes, and failed on a registration that cost another relation more scans than leaving it a view. It left with the register, having nothing left to pair, and what it taught stays: a large pair was a question about the stored table’s index before it was a question about storing it, which is what every large pair it ever pinned turned out to be.

One property of that gate outlives it, and binds any fixture that prices relations. The fixture has to populate the relations it prices: a schema of @table-bound types carrying one scalar field leaves most derived relations empty, and over an empty relation the comparison measures only that H2 charges a table visit at least one scan per naming where a view whose evaluation short-circuits is charged none. That floor is why a small fixture makes such a gate pass while seeing nothing, and why the answer to a gate that costs too much is fewer relations in its domain and never a smaller fixture.

The derivation stratum says what it is doing, and the order it says it in is the whole instrument. It reports to a StageProgress observer the caller supplies: two lines bounding the pass, printed by default, and two lines per stage behind a debug tier, since two lines per stage on every save is not a cadence any default keeps. A stage’s name goes out before its first statement is issued and its duration and row count after its statements return, and that split is not a formatting choice: an instrument that timed each stage and reported afterwards emits nothing at all for the stage that never returns, which is the only case anybody turns it on for. Under the order kept here, a console that stops after the pass line is stuck inside the stratum and one that stops after a stage line is stuck in that stage, so a hang that cost two sessions an afternoon of thread dumps and hand-reproduced sort orders is a name in the first seconds. An observer rather than a logger for a reason that reads as incidental and is not: the property above is a sequence, so a test asserts a list, where a log would need an appender in a module that carries jOOQ and H2 and no logging framework. No threshold and no slow-stage warning, because no honest number exists to set one to: a stage legitimately costs tens of seconds on a large consumer schema, so a threshold that fires on the pathological case fires on every capture there. Ranking the durations is the reader’s job, and a stage getting dearer is a build-time gate’s job rather than a runtime warning’s.

Absence in a derived relation needs a stated meaning, and "not reached" is not "resolves to nothing". Where a chain stops because an element named an unknown key, the elements after it contribute no rows, which is the walk’s own behaviour; where an element carries a condition alone, its destination comes from the condition method’s second parameter, and the parameter resolving to no generated table is a refusal rather than a gap in the walk. Both are silences, and they mean different things, so the view’s comment says which silences it owns. A relation whose absence is load-bearing owes that sentence.

The condition hop is also where the sentence stopped being enough on its own. One no-row there could have meant any of seven things: four typed author errors the resolver states, two silences the classpath census contributes, and the walk’s own "not reached". A comment can distinguish seven meanings but a reader cannot act on one, so the seven became two relations: graphitron_condition_method_route answers where such a hop goes, and a second one named, in a closed vocabulary, why it goes nowhere, which left the hop view’s own absence meaning exactly one thing again. That is the general move whenever a load-bearing absence is asked to carry more than one fact, and it is cheaper than it looks: a relation of reasons stands on the same joins the resolution does.

The instance is instructive for a second reason, and it is why the relation is described here in the past tense. It was built and never wired: no surface read it, no view joined it, and the only thing that ever asked it a question was its own test. It has been deleted. The move is still right and the cost of getting it wrong is visible here, which is that a relation splitting an absence earns its place from the reader that acts on the split, and a split nobody reads leaves the absence exactly as overloaded as it was while costing a relation to maintain. Build the reason relation when the surface that reports it is ready to.

A relation may override rather than answer, and then its silence is most of it. intent_field_column_table says which table a column name written at a field’s site resolves against, but only where that table is not the one the field’s own parent is bound to. Stating the parent’s case too would make it a copy of intent_bound_table keyed one grain down, and every reader of it already holds the parent’s binding. So absence means "the reader’s own default stands", presence means "it does not", and a second disposition says "and nothing stands in its place", which is the case a default would get wrong rather than merely fail to help. An override relation is worth naming as such: its population is not the coordinates a question applies to, it is the coordinates whose answer differs from the default, and a test that pins where no row appears is pinning the boundary rather than reporting a gap.

A silence must not be sourced from a relation that is scheduled to drain. The same view could have read the rejection residue to learn that a coordinate already carries a report, which is a true fact and exactly the reason its incumbent stayed quiet. It does not, because a derivation resting on the residue goes quiet the day that family acquires its own derivation and leaves, and the change would look like nothing: no compile error, no failing pin, just a diagnostic appearing where one used to be suppressed. A derived relation’s meaning must not depend on where a message currently lives, so the silences are structural, and a silence that can only be stated transitionally is better left unstated until the fact it needs exists.

A derived relation is keyed by whatever its own question is about, and the stratum it lives in is decided by what its rows are a function of. The peel established it: what a declared type delivers once its containers are peeled off is a rule over the classpath with no graph in it, so it carried the census’s key, and a graph reached it through store_graph_source like any other source-keyed fact. Keying it by graph would have stored one copy of the answer per graph that reads the class, which is a claim about the graph the rule never makes. It belonged in stratum two all the same, because its rows were a function of captured facts: the transcription families hold what a walk read, and a peel is not something any walk read, it is a computation over the classpath census a walk did. Being the family’s only resident that did not lead with graph_name was the shape of the question, not an exception to it.

That argument has an end, and both of those rules reached it. What member names a class offers an SDL author was the same shape, a rule over the census carrying the census’s key, and it is now code_read_slot: the reading of the classfiles writes it down, because the reading already knows the class’s declared form and has the components in front of it, and a fact the reader of a corpus holds is cheaper stated than recomputed. The peel followed it into code_type_element, and the source-keying is what let it go: a row keyed by its own question rather than by its reader is one the gatherer that owns the corpus can write without asking who will read it. The stratum test did not change and neither did the answer to it; what changed is which corpus the fact is read from. A rule that survives only as a derivation is one no reading is positioned to state.

What a peel drops when it moves is worth stating too, because the move is not free. The derived form was keyed by a position inside a declared type and so could carry that position’s variance, the ? extends or ? super a type argument was written with. code_type_element is keyed by the type and records what the type delivers rather than where the delivery was read from, so variance has nowhere to sit and is not carried. Nothing read it, and the distinction it draws, whether a delivered class can be read out or written in, is the read and write axes the slot relations already state at a level a generator can use. A fact that only one key can hold is a reason to check who needs it, not a reason to keep the key.

A rule that a projection re-runs per build to hand the same answer to several readers is a relation waiting to be written. The bean rule had four readers in the language server and a fifth in the MCP schema resource, and it ran in the catalog builder to project a member list onto each type’s backing shape. Moving it into a view removed the projection’s payload entirely: a backing shape now names a class, and what the class offers is a read. The permits also stopped deciding it, which is what fixed a latent defect in the old shape: because the permit chose which projected list to consult, a record that also declared a bean accessor had two lists that could answer, and the reader took whichever the switch reached first. The relation chooses the arm by the class’s declared form, so there is one answer per name.

Where a macro rewrote a fact, read the written type expression rather than walking the expansion. A connection field’s own named type is the wrapper the @asConnection expansion synthesized, and the columns an author names on that field belong to the element type. The rule that needs the element could walk the expansion’s own fields, which is knowing the shape it expands into; instead it reads the type expression the field was written with and takes its named type. That is the stratum argument, and it is why the derivation is worth stating over the written form: a derivation over it works for any macro that rewrites a type expression, including ones that do not exist yet, while a derivation over the expansion’s shape is coupled to this one. Under the recompute test that also settles which of the two is the captured fact. The written expression is what the walk read, so it is stratum one; the expansion is a function of it, so it is stratum two. The store held the inversion of that for a while, the expansion sitting in graphql_field.type_sdl and the written expression as unparsed text beside it, which is why two views once recovered it with nested REPLACE calls stripping [, ] and !. Both are ordinary columns now: graphql_field holds the author’s expression with its wrappers already decomposed, graphitron_field_minted holds the macro’s replacement as a row whose coining coordinate is the field’s own, and graphitron_field is the two resolved. What made the correction hard was that this expansion rewrites a captured value rather than adding a row beside one, so no anti-join recovered the author’s side; the answer was to stop overwriting rather than to recover better. The rule the paragraph above states is unchanged and is now a plain join: a derivation wanting what the author wrote reads the transcription.

Violations are located facts, not log lines. A violation is a row, not an act: the diagnostics stratum records rejection, lint, advisory, SDL-toolchain and compile rows (the rejection_, lint_, build_warning_ and javac_ families, plus the two verdict residents of graphql_), and the deliberately prefix-less diagnostic view unions them into one read surface with the rendered coordinate computed from the stored columns. Keeping the violation a relation makes every surface a projection of it (the LSP diagnostic, the MCP triage tools, a build report); baking it into one consumer’s output would force every other consumer to recompute the rule.

Gathering the schema is itself a staged pipeline, and a stage’s refusal never cancels the next stage. Five stages, each with its own unit:

Stage Unit Owns

1. Per-file parse

one file

source membership, graphql_schema_problem at stage PARSE, and the per-site declaration facts, which carry the file’s own cadence and therefore survive any later stage’s refusal

2. Combined registry

the file set

the declarations admitted together, and graphql_schema_problem at stage REGISTRY for what the registry refused

3. Composition and assembly

the whole schema

the registry composed with the loading rewrites the generator runs and assembled into a GraphQLSchema, graphql_schema_problem at stage REWRITE where a rewrite refused and at stage ASSEMBLY where it did not assemble

4. Coordinate-keyed census

the file set

the merged effective element set: one row per coordinate, both occurrences of a duplicated one already rows of the entry stratum

5. Rooted traversal

the assembled schema

intent_type_domain, the classification domain’s type members, from the seeds the document states

The first four judge, and the three that judge a document write their verdicts down. The parser judges one file at a time, then the registry judges the combined declarations, then the loading rewrites compose the registry and assembly judges the composition against the GraphQL specification’s structural rules; graphql_schema_problem holds every refusal, the stage a column, because parsing a file, merging the files and building a schema of the result are not separate questions a reader has. A rewrite that refuses never stops the stage: the assembly judges the merged registry instead, so what it records after the refusal is a true statement about the corpus as written. Each stage keeps what survived it, so a source that will not parse costs its own declarations and no others, and a declaration the registry will not admit costs itself. That is what makes assembly worth running unconditionally, whether or not a pass has any use for the assembled schema: it is the only place the specification’s structural rules get checked at all, so its verdict is a fact about the consumer’s schema worth as much as the declarations it judges. Aborting at the first refusal instead is what lets one freshly broken file blank every fact about every file beside it, which is precisely when an author needs those facts most.

Assembly is the gatherer’s own stage, not something a caller hands it a verdict about. Stage 5 reads the schema stage 3 produced, so taking the transcription and the verdict from one assembly is what keeps the store from judging a document it does not hold. It also decides which registry gets assembled: the corpus the store transcribes, composed with the loading rewrites through the function the generator’s load calls, and before the synthesis rewrites that inject declarations. The generator’s own verdict assembles the same composition, so a directive the build accepts cannot be reported undeclared here. Judging the post-synthesis registry instead let a verdict blame the author for a declaration graphitron’s own rewrite added. The decode does not walk the composition: it walks the merged corpus the composition started from, a reduce of its own, so the as-written half of the store is a function of the documents alone.

Composed versus written is decided per relation, never per stage. The per-site declaration facts are stage 1’s and stay verbatim at their source’s cadence. What composition adds over them is either a view (an argument default is deliberately one join away, and filling it in at capture would store a derivation where a view answers) or its own coordinate-keyed relation whose comment owns why it is not a view. Stage 5 is the case that earned a relation: the closure over a cyclic type graph has no safe H2 view form, and the descent rule is graphql-java’s own child semantics rather than something SQL should restate edge kind by edge kind. Only what genuinely cannot exist without an assembled schema sits behind the assembly stage, which is why intent_type_domain is empty on a run whose registry did not assemble while every declaration fact beside it is current: absence there is read together with the ASSEMBLY verdict, never alone.

The reference web anchors on existence, never on attributes. Each SDL coordinate has a graphql_*coordinate relation carrying its key and nothing else, every foreign key in the schema that names an SDL coordinate names one of those, and the attribute relations beside them (graphql_type, graphql_field, graphql_argument, graphql_enum_value) are referenced by nothing at all. The reason is cadence and the failure it prevents is concrete. A coordinate’s existence is settled by the per-file parse, while an attribute may be owned by a later stage that can be withheld, rewritten or refused, so a foreign key across that boundary makes every per-site fact’s existence conditional on a stage it owes nothing to: one dangling type reference in a half-typed file would take the whole graphitron decode family with it, which is the failure above in its worst form. Before the split, graphql_type asserted both that a name exists and what its kind and description are, and the decode families and site rows had no choice but to anchor on the pair; a stage wanting the attribute half could not take it without taking the family’s availability too. Reading it: join the coordinate to ask what exists, join the attribute relation to ask what it is, and expect the two populations to agree wherever no later stage has taken an attribute over. One direction of that agreement is a foreign key (an attribute row with no anchor cannot be written); the other has no constraint that could see it, so FactCaptureAgreementTest pins both.

The four anchors and what hangs off each:

Element Anchor relation Attributes and dependents keyed on the anchor

type

graphql_type_element

graphql_type (kind and description), graphql_type_declaration (every declaration site), graphql_poly_member, the type-keyed graphitron_ decodes, intent_type_domain, intent_type_backing_class

field

graphql_field_element

graphql_field (type expression, description, default, contributing site), the field-keyed graphitron_ decodes, intent_input_occurrence_path_step

argument

graphql_argument_element

graphql_argument, the argument-keyed graphitron_ decodes, intent_input_occurrence_path

enum value

graphql_enum_value_element

graphql_enum_value, graphitron_ast_enum_value_binding_entry

The directive-definition family has no anchor of its own and needs none: a directive is defined once, nothing merges into a definition, and only its own two children reference it.

What an author applied is not in the table above, and that is the point of it. graphql_directive_application keys on the coordinate rather than on any of the four anchors, so one relation answers at every site including the schema block, which is not an element of the specification’s grammar and carries the coordinate $schema for this purpose. It was five relations, one per site, and they differed only in which decomposed key they copied down; the coordinate is what all five had and none of them stored. Its arguments hang off it in graphql_directive_application_arg, which was five relations for the same reason and is now one.

Enforced by: SdlCoordinateCensusTest pins capture’s merge against graphql-java’s own composition at all four grains, and every merge-ordered ordinal family by value rather than by density, on a fixture whose base definitions and extensions are deliberately out of order. The ordinal families are merge-ordered because capture numbers a type’s fields, arguments, enum values, union members and repeated directive applications with counters it holds per type and carries across the declaration sites, so a family checked only for density passes with the merge order inverted. FactCaptureAgreementTest pins each anchor against its attribute relation in both directions, the direction no foreign key can see.

Freshness is not a property any consumer carries. It used to be carried once, on the snapshot handle a consumer held, along two axes: whether a pass had produced a projection at all, and whether the one in hand was the latest parse or the last good one before a regression. The consumer split was purely which precondition each read imposed, and the editor’s was to tolerate the previous projection and tag what it showed, rather than punish an author for a half-typed edit. Two-stage capture removed the state that axis described. A file the parser refuses and a schema the assembler refuses are both rows, so there is no moment at which the store withholds an answer and nothing for a consumer to tag; the axis retired with its last reader. What is left is a lag of stated size rather than a state to switch on: the store answers for a file’s last saved content, one save behind the buffer, and a surface needing the unsaved text reads the buffer for it directly. The referenced namespaces (the jOOQ catalog, the classpath census) still lag their sources, and that lag is now visible as each family’s own capture cadence rather than as an axis on a handle.

Location is a fact about an entity, and the rule is cadence rather than storage. A position stored on a relation that refreshes more slowly than the position does freezes stale, which is why Java positions and Javadoc are not columns on code_class: the classfile census refreshes when a build produces classes, and a declaration moves when somebody saves a file. What follows from that is not that the position cannot be stored, but that it is stored on its own cadence. The java_ family holds one row per declaration a source parse read, keyed by the file it is written in and refreshed per file when that file changes, and a consumer joins it by name to the census, or to a captured generated-class FQN for the jOOQ half the census deliberately excludes. The relation is partial by design (a built-in has no position, and a headless session that walks no sources has none of these rows at all; a read allowed to be empty must not be a key), and the join is outer on both sides, because the two populations answer as of different moments and a view asserting they agree would assert something neither walk knows. SDL positions were always the same rule’s sanctioned side: the SDL is what capture parses, so its positions arrive in the same pass as the facts they locate and are stored as capture columns (source_name, source_line, source_column), in the declaration-site key even, because two extensions of one type can share a line. "Joined, not stored" is the law for positions that move on a cadence the fact does not; the principle underneath it is that a fact refreshes on the cadence of its own source.

Enforced by: DiagnosticFactsTest pins the diagnostics stratum’s derived columns against their Java spellings; DiagnosticsAggregateTest pins the union view’s shape through the shipped MCP tool; FactCaptureAgreementTest’s oracle-lifecycle gates pin that each post-capture writer clears exactly its own graph partition and nothing else; `JavaSourceFactsTest pins the source cadence’s own ownership scope, a file at a time, and CatalogRefreshTest pins that a .java edit moves the store row with no generator round. The materialized-view prohibition is enforced by the build itself rather than by a test: graphitron-model’s jOOQ codegen reads `INFORMATION_SCHEMA.COLUMNS off a live store booted from this schema, so a CREATE MATERIALIZED VIEW added anywhere in graphitron-model.sql fails the model build before any test runs. The LOCAL TEMPORARY requirement has no enforcer and is a trap to read rather than an invariant to lean on.

One base, many views

Every consumer (code generation, the language server, the MCP knowledge surface, the test corpus) reads views over the one base; no consumer owns a private model. Code generation is the narrowest view, not the model: it reads the resolved values and demands a total, integrity-clean snapshot, and it simply does not project the columns the other consumers live on (the authored form behind each resolved value, the description text, the source position). Those columns are in the base regardless.

The invariant that keeps it one model through the migration: every consumer re-sources onto the store, and each view’s coverage guarantee moves with its projection seam. A migration that leaves one consumer reading the old surface revives the leaves as a shim purely to feed that consumer, and the model forks; facts, revived leaves, and projections is three models. The strangler frame prevents it: consumers migrate one at a time, new facts land only in the store while both models are live, and the two-model window shrinks monotonically instead of fossilizing.

Enforced by: FactCaptureAgreementTest keeps the two live pictures honest for every relation while the window is open. No projection seam is left to gate: the classification projection was the language server’s view of the classifier, and it deleted with its last reader, so code generation classifies into its own model and reads that directly. Each consumer’s coverage gate is therefore over its own sealed vocabulary rather than over a switch two consumers share. GeneratorCoverageTest.everyGraphitronFieldLeafHasAKnownDispatchStatus partitions every model leaf by dispatch status, so a new variant is generated, declared unsupported, or breaks the build; TriggerDispatchMatrixTest pins the language server’s request triggers against its surfaces as an exhaustive partition, so a new trigger is answered, declared unanswered, or breaks the build.

A consumer’s answer is one projection at its own grain, never several grains folded back together in the consumer. jOOQ’s MULTISET nests a one-to-many child under its parent through the foreign key the child relation already declares, and Records.mapping lands each level on the record it already has; H2 serves the nesting by emulating it over JSON aggregation, two levels deep and with row(…​) inside a multiset. The alternative, several statements at several grains reassembled with accumulators and a synthetic grouping key, is a relational join written in Java, and it fails in three ways the nesting cannot. The grouping key has to be invented and can be invented wrong: a constraint name is unique per table and not per schema, so folding both foreign-key directions needs a four-part key that the relation states and the Java has to remember. Consistency has to be argued rather than held, since several statements need a read transaction wrapped round them where one statement is atomic. And the row count crossing the JDBC boundary is the product rather than the sum, a parent row repeating once per child. More than one statement is still right where the questions are genuinely independent, a page beside its unpaged total, or where two relations share neither a key nor a partition dimension and the store declines the correlation on purpose, as the classfile census and the source-declaration family do. What the rule forbids is one answer assembled from several grains after the fetch.

Picking the grain and picking the relation the statement drives from is one decision, taken before any SQL is written. Name what one row of the answer means, and the natural key of that sentence is the answer’s grain; the relation owning that key goes first in the FROM clause, and the rest attaches to it through keys the relations already declare. A child grain is never folded into that projection: it nests as a correlated MULTISET on the key its own relation declares, or it becomes a second statement paired on a real key, which is the shape SchemaQueries reads the MCP schema surface at, two statements at two grains paired on the type’s own key. Grain is stated prose rather than inference, so the register is the same one the store already keeps: a relation’s COMMENT ON says what one of its rows is, and an answer’s grain is a sentence of that kind about the projection. The smell is the reverse order, driving from whatever the caller happened to hold and repairing the shape afterwards with a grouping key the store never stated or a DISTINCT that hides a fan-out instead of answering it. Where the grain’s owner is a derivation carrying a window function or a recursive term rather than a base relation, the rule under "Derived reads are views, not stored facts" decides which way round the statement goes, and it is the view’s own shape that decides, not the reader’s spelling.

An empty relation is a fact about the population, and a reader that wants to distinguish "this name is wrong" from "nothing has told us about these names yet" must get both answers from one read. That distinction is real: a consumer who has not run their build yet has a census with nothing in it, and a surface that treats every name as unresolved then is wrong about every one of them. Splitting it into a presence check before a lookup makes correctness depend on the order a caller happens to write them in, and each caller writes them again. So the reader answers with the arms, not with a boolean: known, unknown, nothing captured. The shape follows what an arm carries; an arm set where none of them carries anything beyond its identity is an enum, and one where a single arm carries rows is a sealed interface. Where several arms of a surface share the deferral, it is decided once for all of them, at the dispatch, rather than restated per arm.

The back half: complete commands, a closed graph

The emit side consumes the facts as commands, and the law is that commands must be complete: the render shell makes no decision the planner could have made. A command row carries everything its renderer needs, and the command / plan / render package triangle keeps producers and consumers apart, with the emit library visible only to render.

The emit target is a referentially-closed graph of Java methods: a node relation (the methods the plan committed) and an edge relation (calls by name). Closure is bidirectional. Every method the generator emits is the render of exactly one committed command, and every callee name in every emitted body resolves to a method the run also emitted; a renderer minting a callee name the plan never committed is exactly the leak the invariant exists to catch.

Two rules govern the graph’s shape. The seam-placement rule: a named method call (a seam) belongs where a unit is chosen by a runtime dispatch, reused across more than one caller, or something the tests must assert independently; inline only a linear, single-use, non-varying construction. The single-mint naming regime: a callee name is minted once, in the plan’s naming vocabulary (GeneratedUnits), and read blind on both ends of the edge, never reconstructed by formula at a call site.

Enforced by: PackageImportDirectionTest for the triangle, and its unitRefsAreMintedOnlyByThePlansNamingVocabulary for the mint (unit references are constructed only inside GeneratedUnits, checked over the whole main tree); MethodClosureOracleTest for callee resolution over a seam-spanning schema; LauncherRelationClosureTest for the bidirectional half over the launcher relation (every covered coordinate has exactly one row, every row resolves to an emitted method, no two rows claim the same method).