Inside Tech Comm with Zohra Mutabanna

S8E8 The Content Architecture Beneath AI with Patrick Bosek

• Zohra Mutabanna

Use Left/Right to seek, Home/End to jump to start or end. Hold shift to jump forward or backward.

0:00 | 50:07

AI can retrieve and generate answers, but the quality of those answers still depends on the information beneath the model. If content is poorly structured, weakly governed, or divided into the wrong units, a more capable LLM does not automatically solve the problem.

Patrick Bosek, co-founder and CEO of Heretto, joins me to examine how structured content shapes AI context. We begin by defining structured content and comparing the different approaches, from DITA and XML to headless content management systems. Patrick explains why the right structure depends on scale, interoperability, reuse, and the ways an organization needs to deliver information.

This conversation builds on episodes Documentation Is Becoming the Interface and The Work AI Cannot See. If documentation is becoming a surface through which people interact with products—and AI cannot infer context it has never received—then content architecture becomes part of the AI system itself. This episode moves one layer deeper to examine that foundation.

Patrick walks through chunking, retrieval-augmented generation (RAG), vector search, metadata, and knowledge graphs without treating AI as a replacement for established information practices. Metadata still provides deterministic controls. Traditional search may remain the better choice in some situations. Governance and provenance still matter when content comes from product, engineering, support, and documentation teams.

We also consider the practical side. A lone writer can begin with a small style guide, consistent Markdown, useful metadata, and an AI agent that reviews content against established rules. At enterprise scale, however, tools alone are not enough. Information modeling, information architecture, governance, and deliberate implementation have to come first.

In this episode, we discuss:

  • What structured content means beyond a single format or tool
  • When DITA becomes useful—and when it may be more than a team needs
  • Why clean content boundaries improve chunking and retrieval
  • How RAG systems retrieve information outside the language model
  • Why metadata remains valuable in both deterministic and AI-driven systems
  • When traditional search can compete with or outperform vector search
  • How metadata can connect structured content with knowledge graphs
  • Establishing authority, provenance, and a source of truth across teams
  • Practical governance steps for lone writers working in Markdown and GitHub
  • Why implementing a CCMS also requires information modeling and architecture
  • Why a well-structured website remains important when public AI systems answer customer questions
  • The difference between buying tools and building the information foundation those tools require

Season 8 of Inside Tech Comm is sponsored by LavaCon. Use discount code ITC26 to save $200 on registration for LavaCon 2026.

Show Credits

  • Intro and outro music - Az
  • Audio engineer - RJ Basilio