Obscure RDF formats

Opinion
Semantic Web
RIF, GRDDL, POWDER, HDT and other obscure W3C formats: standards are easy to make and hard to get adopted. Why HDT is cited often but used rarely.
Published

September 1, 2026

Illustration: Obscure RDF formats

Obscure W3C RDF-formats are all around us: RIF, GRDDL, POWDER, SAWSDL, XForms, OWL FULL, CURIE and many more. It’s one thing to make a ‘standard’ but whether people adopt it is something else. Same for HDT (Header, Dictionary, Triples), it gets cited constantly in RDF compression papers. It almost never gets cited in enterprise architectures.

HDT is a binary RDF serialization that separates a graph into three parts: - the Header: plain-text metadata, provenance, statistics, VoID-style descriptions - a Dictionary: every IRI, blank node, and literal gets a numeric Id, sorted into Shared/Subjects/Objects/Predicates for locality - Triples: BitmapTriples encode the graph as a forest of trees rooted at each subject and referencing dictionary IDs instead of strings.

The real result is an order-of-magnitude smaller than Turtle or N-Triples, it’s small enough to sit in memory and triple-pattern lookups work directly on the compressed form. Cool. Well, not really.

The business case gets thin in various ways. First off, HDT is read-only. No updates, no deletes, no inserts. For anything that changes more than occasionally that’s not a trade-off, it’s a disqualifier. Most enterprise knowledge graphs I see are not static. Second, the tooling ecosystem is small and mostly academic. A handful of C++/Java libraries, a reference spec still sitting as a W3C Member Submission rather than a full standard. Third, it solves a storage-and-transfer problem, not a reasoning or integration problem. Compressing your graph doesn’t help you map it to an ontology, keep it consistent or connect it to an LLM’s context window. Finally, almost none of the current GraphRAG stack touches RDF or HDT at all. LangChain, LlamaIndex, Microsoft GraphRAG, Graphiti… all property graphs and vector stores, not triples. Pitching HDT into that conversation is usually a mismatch.

None of this means the underlying ideas are wrong. A handful of the standards did stick: Turtle, SPARQL, SHACL, JSON-LD earned adoption because they lowered the cost of doing something people already needed to do. At the same time, the RDF graveyard is bigger than the survivors and sometimes I feel very concerned when I tell customers about possible semantic solutions. A spec, even a W3C spec, is not evidence anyone is or will use it. Ask who has shipped it under load… Is the imbalance a maturity problem the field will grow out of or a structural mismatch between how W3C standards get made and how enterprises adopt tools? Whenever I hint at this the RDF guards are quick to send me more research papers and unconvincing singular business cases. Reminds me of the R statistical language which also grew out of academia and still feels awkward as an enterprise tool. The proverbial wisdom that the road to hell is paved with good intentions comes to mind.