The universal architecture for knowledge augmented AI

Many people and companies ask me for the same architecture, even if their businesses differ widely it always boils down to the same I-want-to-have: provenance, exposure, bitemporal, reasoning, graph RAG, ontology, data governance and more. I usually listen with patience to understand where they stand in graph-world and in which direction to point. Sometimes it feels like I am a touristic guide in the confusing land of (knowledge) graphs. A reference architecture is missing and I should take time to assemble one. The data size would be a crucial aspect (anything works over a Gb of data, even JSON). Another thing is my conviction that you need both RDF and LPG to make things happen. Not a popular opinion and one I often can’t convey because it comes with a steep investment for the customer in terms of money, time and dev learning curve. Long story short, a recent benchmark called “CypherBench” uses Wikidata with a projection to Neo4j and instead of pointing an LLM at the raw RDF they generate materialized views on top of it: schema-scoped, denormalized, Cypher-queryable slices of the same underlying data. Clean node types, readable edge labels, no identifier noise. They did this across 11 domains, 7.8 million entities, and built over 10,000 (!) natural-language-to-Cypher questions to test it. I wished I had done this.
The result isn’t ‘RDF loses to LPG’ or something, It’s a division of labor. RDF+SPARQL keeps provenance, governance, T-Box reasoning and is the layer where correctness is non-negotiable. A materialized LPG view becomes the layer the LLM actually talks to. Scoped to a task, denormalized for readability, disposable and regenerable.
This is the same KAAI pattern (knowledge augmented AI) I keep landing on with clients asking for heaven on Earth: don’t ask a reasoning-grade store to also be an LLM-friendly retrieval surface. Let the triple store keep its rigor and project a clean view for the layer that needs to be legible to a model. Throw the view away and regenerate it when the schema moves.
The graph representation is in this setup the durable decision. The interface you expose to a model on top of it is not and treating it as disposable is what makes RDF and LPG work together instead of competing.
CyberBench: https://github.com/megagonlabs/cypherbench