The generator turns .graphqls files plus a jOOQ-generated database catalog into Java sources. The architecture is fact-oriented: a run transcribes everything it reads into a relational fact store, derives claims and violations as views over those base facts, joins facts into complete command rows, and folds a render shell over the committed commands. During the current strangler window a classification walk still produces the model most consumers read; that walk is a transitional surface, named as such below, and must not be mistaken for the destination. This page owns the stages and their order; the modeling discipline behind the store is The fact model.

flowchart LR
    A[".graphqls files"] --> B["SchemaLoader<br/>(parse + auto-inject<br/>directives.graphqls)"]
    B --> C["GraphitronSchemaBuilder<br/>(classification walk;<br/>transitional surface)"]
    B --> D["FactCapture<br/>(total transcription into<br/>the store's base relations)"]
    D --> E["Claims and violations<br/>(intent_ claim views,<br/>violation facts)"]
    C --> F["GraphitronSchemaValidator<br/>(fused with store-derived<br/>violations)"]
    E --> F
    F --> G["EmitPlan<br/>(join facts into<br/>command relations)"]
    G --> H["Render shell<br/>(fold over committed<br/>commands)"]
    H --> I["JavaFile.writeToPath<br/>(idempotent writes,<br/>orphan sweep)"]
    I --> J["consumer compile<br/>(graphitron-sakila-example)"]

The stages below appear in the order GraphQLRewriteGenerator.runPipeline runs them. Each stage has a verb: capture transcribes, classification gathers, validation derives, planning joins, the shell folds.

Parse and attribute

SchemaLoader merges the consumer’s .graphqls files and auto-injects directives.graphqls from the graphitron-model jar before parse, so consumer schemas never re-declare the canonical directives and every later stage can treat directive presence as ground truth. The attribution pipeline then applies the schema-level rewrites (federation @link handling, tag and description notes) and cuts a read-only pre-synthesis snapshot of the registry before KeyNodeSynthesiser rewrites federation keys. Capture reads that snapshot and runs the @asConnection expansion itself rather than inheriting the rewrite’s output, but not inside the SDL walk: the expansion is the last step of the gatherer that decodes the documents, anchoring out of the graphitron_connection_entry rows that same gatherer has just settled. Nothing it produces lands in the transcription. What it mints is graphitron_type_minted, graphitron_field_minted and graphitron_argument_minted, which are views over the applications rather than tables the expansion fills, with graphitron_minted_coinage naming which coordinate’s application coined each minted type; what it rewrites is a graphitron_field_minted row at the rewritten field’s own coordinate, flagged as a rewrite, holding the replacement expression while the author’s stays in graphql_field where it was written. So both readings are plain rows at the same coordinate rather than one value and a provenance record no anti-join reaches, and graphitron_type and graphitron_field are where a reader that wants the population the generator emits finds the two together. Federation’s key synthesis is not among the expansions capture runs: its rule reads node-identity metadata off the generated jOOQ classes as well as the SDL, so it is a derivation over the captured facts of both corpora (graphitron_synthesized_federation_key) rather than something the SDL crawler may decide.

Classification gathers (transitional)

GraphitronSchemaBuilder walks the attributed registry once and classifies every type and field into the sealed GraphitronSchema model, gathering per-trigger facts as it goes. This walk and the leaf model it produces are the transitional producer surface of the strangler migration: they stay live until each consumer re-sources onto the store, and they are a surface being drained, not a place to extend. New facts land only in the store; a new capability is added by adding a fact relation, never a new leaf type or walk-side registry. The zoomed-in view of this surface, the classification taxonomy and what each verdict drives, is Code Generation Triggers.

Capture transcribes: the fact store

ModelCapture is the pass, run by whoever owns the store: a mojo or a test opens one and hands it over, and nothing in the generator can open one. It transcribes what the run read in one transaction, the parsed SDL into the graphql_ relations, the jOOQ catalog into sql_, and the classpath into code_, with store_ bookkeeping recording the graph, its sources and what their bytes hashed to. Each corpus is read once by the gatherer that owns the store’s record of it, and the gatherers that write from it are handed that reading. The graphitron_ relations fill in the same transaction: decoding the graphitron and federation directives is a derivation over the applications graphql_ just captured, read back out of the store rather than handed over while a walk still holds the parse. The gatherers run in the order their declared read edges require (meta_gatherer_dependency), and each one’s rows reach the store before the next starts: a flush is not a commit, so a gatherer reads what ran before it through the store, and nothing outside the transaction sees a partition mid-load. Capture is total and tolerant, transcribing what is there including shapes later stages reject, so the store is a faithful record rather than a filtered one. A source whose bytes have not moved since this graph last transcribed them is left alone, compared against the claim the reader wrote; a persisted store lives under target/ keyed by DDL hash and generator version, so an incompatible upgrade opens a fresh file instead of migrating, and mvn clean removes it. Once every gatherer has flushed, the graphitron gatherer’s derivation stratum runs: the stages DerivationStratum lists, in the one order they may run, each clearing its own partition and writing it again, most of them one INSERT over the stored _rule view that states its rule and a few hand-written producers, InputOccurrencePaths and the classification domain (ClassificationDomainCapture) among them, where H2 cannot state the rule as a safe view. On a warm store the stratum runs inside the capture transaction; on a store none of whose stage tables holds a row it commits and analyses stage by stage, so every stage plans against statistics for what the stages before it wrote.

Claims and violations as facts

Above the base relations the DDL defines the intent_ stratum as views: the authored claim views (intent_authored_field_claim, intent_authored_type_claim) union one arm per claiming relation, the resolved views reduce them, and intent_authored_claim_conflict detects coordinates carrying contradictory claims. AuthoredClaimConflicts reads that view and projects each conflict row into the same located ValidationError the retired walk-side detector sites used to produce. The diagnostics stratum (rejection_, lint_, build_warning_, javac_ relations) records rejection, lint, advisory and compile rows, unioned by the deliberately prefix-less diagnostic read surface that the MCP diagnostics tools serve; why a violation is a row rather than a log line is The fact model. The diagnostics writers run in the dev session, which holds a live store handle across builds.

Validation derives

GraphitronSchemaValidator derives the run’s verdict from the classified model, fused with the store-derived violations from the conflict detection. Rejection is a typed value end to end, never parsed prose; the taxonomy and its contract are Typed rejection. A failing verdict throws before any file is written.

Planning joins

EmitPlan.produce joins the model’s facts into the command relations the run will render: the launcher relation (one row per covered coordinate), the condition relation (one row per filtered coordinate and resolved table), the projection relation (one row per projection unit), the fetcher-edge relation, the type-unit relation, and the global commands. Commands are complete rows; the completeness law, the package triangle that enforces the producer/consumer split, and the naming regime are The fact model.

This layer is a functional core with an imperative shell, and two rules follow from that. A planner holds its own queries, decides what the run emits, and returns a command relation; the relation is immutable data, so it asks the store nothing and carries no handle to ask with. And what a planner needs is its own business, so two planners wanting the same rows each write their own query rather than sharing one. A derivation both would want is a fact missing from the model rather than a helper missing from this package, and the fix is to push it down.

RoutineWriteCommands does not have this shape. It takes the handle and reads through RoutineWriteFacts, a query surface built for reuse. That is the older arrangement rather than the one to copy, and it dissolves as the arcs owning those commands reach them.

Render: the shell folds

GraphQLRewriteGenerator.runPipeline folds over each committed relation and hands every row to its renderer (RootLauncherRenderer, ConditionGlueRenderer, ProjectionUnitRenderer, and the schema, record and global emitters). The fold enforces closure in both directions: a renderer emitting a unit the plan never committed fails the run, and a committed unit no renderer emitted fails it too. The closure invariant over the emitted method call graph, and the oracles that pin it, are The fact model.

Write: the idempotency contract (unchanged)

On every run, JavaFile.writeToPath writes only files whose rendered content differs from disk (SHA-256 comparison) and the generator deletes orphans in rewrite-owned sub-packages. Both halves run on every emit, not just full builds; this is what keeps the dev-loop’s IDE-recompile times proportional and what stops a delete-a-type cycle from leaving stale files behind. Pinned by IdempotentWriterTest and GeneratorDeterminismTest.

Consumer compile

graphitron-sakila-example compiles the emitted sources with <release>17</release> and runs the execution tier against a real PostgreSQL, closing the loop: the pipeline’s output is verified as compiling, type-correct, behaviorally-tested Java, not just rendered text.

The strangler frame

The store shipped beside the live walk, not instead of it, and the migration drains one consumer at a time. What that means concretely today:

  • Capture is total, but most relations are populated ahead of their readers. The store’s production readers today are the authored-claim conflict detection (whose verdicts also feed the LSP/MCP conflict overlay) and the diagnostic view; planning still joins facts gathered by the classification walk. FactCaptureAgreementTest keeps the two pictures honest with a mechanical driver over every generated relation and no skip list, so a new relation cannot arrive unchecked.

  • GraphitronSchema and the leaf classifier stay live until each consumer re-sources onto the store. Extending them is migration debt; the rule during the window is that new facts land only in the store, per the re-sourcing invariant in The fact model.

  • The schema gates (FactSchemaGateTest) hold the store’s own invariants: total comment coverage, dense ordinals, verbatim-transcription twins for every decode, and the graph-partition rules. A failure there is a capture bug or a DDL defect, never an author error.