📁 ID_0003 · Portfolio Artifact
Allana Jackson · Knowledge Systems Architect (Contract)
DoorDash IT · Knowledge Infrastructure Transformation · 2024
Knowledge Architecture Semantic Search · ML-informed Retrieval Confluence · GitHub · Python
Internal Design Document · IT Knowledge Management · v2.0 · Final
Knowledge Infrastructure Transformation Blueprint
Audit findings, information architecture design, semantic search parameter specification, and governance model for consolidating DoorDash IT's distributed documentation ecosystem into a unified, searchable knowledge platform.
Confluence · GitHub · Jira · Markdown Okta / SAML · Auth Flows · SDK Documentation 25+ Repositories Consolidated ML-informed Semantic Search · Python · Runbook Templates
25+
Fragmented
repositories
→
 
1
Centralized
platform
75%
Faster
retrieval
📋 Discovery Methodology — Knowledge Governance Architecture

I architected a three-phase knowledge audit system to map the full scope of documentation debt across DoorDash IT engineering — treating the investigation itself as a structured design process, not a survey. (1) Repository discovery: I built the inventory framework and catalogued all 25+ active and abandoned documentation sources by team, tool, last-modified date, and ownership status, cross-referencing Jira incident tickets to surface undocumented operational gaps. Sources spanned Confluence, GitHub repositories, Notion, Google Drive, and Markdown-based sources. (2) Stakeholder interviews: I designed and ran structured discovery sessions with engineering, platform, security, and on-call teams to extract undocumented pain points and quantify friction against a consistent framework. (3) Pain point classification: I developed a friction-impact matrix to tier findings by severity and sequence the consolidation and remediation roadmap — forming the foundation of the Knowledge Governance model rolled out in Phase 5.

Repository Inventory — Pre-Consolidation State
25+ active documentation sources across 6 teams · Border color = severity of fragmentation pain
BEFORE STATE
📄 Critical
Incident Runbooks (Legacy)
incident-response / Confluence
📁 47 pages ⚠️ 60% stale
🔐 Critical
Auth / SSO Guides
security-eng / GitHub
📁 12 pages ⚠️ No owner
🏗️ High
Service Architecture Docs
platform-eng / Confluence
📁 31 pages ⚠️ Fragmented
🚀 High
Onboarding Materials
eng-enablement / Notion
📁 22 pages ⚠️ 3 versions
🔌 High
API Integration Guides
platform-eng / GitHub
📁 18 pages ⚠️ Outdated
📊 Medium
Monitoring Playbooks
sre-team / Confluence
📁 14 pages ⚠️ Incomplete
🛠️ Medium
Deployment Procedures
devops / GitHub
📁 9 pages ⚠️ Siloed
📋 Medium
Change Management Logs
it-ops / Google Drive
📁 38 docs ⚠️ Unstructured
🧪 Medium
SIT/UAT Test Plans
qa-eng / Confluence
📁 16 pages ⚠️ Mixed formats
📚 Low
Product Specifications
product / Confluence
📁 27 pages ✓ Current
🌐 Low
Network Topology Docs
infra-eng / Markdown
📁 8 pages ✓ Maintained
⚡ Medium
+14 Additional Sources
various teams / mixed tools
📁 90+ docs ⚠️ Audit pending
Interview Findings — Pain Point Matrix
Structured interviews with 6 engineering teams · Friction scored 1–10 · Impact scored 1–10
Team Primary Pain Point Content Gap Friction Score Eng. Hours Lost/Wk Priority
Incident Response Runbooks scattered across 4 tools; on-call engineers can't locate procedures under pressure No single source of truth for escalation paths or rollback steps
9.5
~8 hrs P0
Security Engineering Okta/SAML guides undocumented or sandbox-untested; credential incidents tied to missing auth documentation No sandbox-validated auth integration guides; no credential rotation procedures
9.0
~6 hrs P0
Platform Engineering Service architecture docs outdated or missing for 40% of services; new engineers re-investigate solved problems No "Golden Path" onboarding; no service catalog with ownership data
8.0
~5 hrs P1
SRE / DevOps Monitoring playbooks and deployment procedures exist in at least 3 locations with conflicting versions No versioning strategy; no deprecation process; no review cadence
7.5
~4 hrs P1
Engineering Enablement New-hire onboarding lacks a structured path; 3 competing "getting started" guides with no official designation No canonical Golden Path; inconsistent tooling setup instructions
7.0
~3 hrs P2
IT Operations Change management documentation unstructured in Google Drive; no search, no standards, no templates No modular templates; no approval workflow documentation
6.0
~2 hrs P2
Synthesized Findings — 4 Core Themes
Cross-team patterns extracted from interview data and repository analysis
01
Retrieval Failure Under Pressure
Engineers consistently failed to locate the correct runbook within acceptable time during live incidents. Average retrieval time was 12 minutes across teams — four times the 3-minute target established as the platform goal.
"I know the runbook exists somewhere. I just can't find it when the alert is firing." — On-call Engineer, Incident Response
02
Tribal Knowledge Dependency
Critical operational knowledge lived in Slack threads, individual engineers' notes, and informal channels rather than in any documented system. Engineer attrition represented an active knowledge-loss risk.
"If you want to know how that auth flow works, you have to ask Marcus. It's not written down anywhere." — IT Operations Lead
03
No Single Taxonomy
Each team had independently developed naming conventions, folder structures, and content formats. The same concept (e.g., "service restart procedure") appeared under 6+ different titles across 4 tools — with no cross-referencing.
"I wasn't sure if I was looking at the latest version or something from 2 years ago. The dates were wrong in half the docs." — Platform Engineer
04
No Ownership Model
78% of audited documents had no designated owner, no review date, and no staleness indicator. Documentation was treated as a one-time artifact rather than a living product requiring maintenance.
"Nobody knows who owns half these Confluence pages. They just exist." — Senior SRE
State Transition — Information Architecture & Content Strategy
From fragmented, unsearchable silos to a centralized, semantically indexed knowledge platform — designed through a structured IA and content strategy process
IA · CONTENT STRATEGY
✗ Before — Fragmented State
❌25+ disconnected repositories across Confluence, GitHub, Jira, Notion, Google Drive, and Markdown files
❌No shared taxonomy — each team used independent naming, structure, and formats
❌No search across sources — engineers used Slack and tribal knowledge to locate docs
❌No content ownership model — 78% of documents had no designated owner or review date
❌Duplicate and conflicting versions — same procedures documented in 3–4 locations with no canonical source
❌Critical auth guides (Okta/SAML) untested in sandbox — driving credential-related production incidents
❌Average retrieval time: 12 minutes under incident conditions
→
✓ After — Unified Platform
✅Single centralized Confluence platform with standardized space architecture, GitHub integration, and Jira-linked incident tracking
✅Unified taxonomy — 4 content types, consistent naming conventions, cross-team hierarchy
✅Semantic search layer with engineered parameters enabling concept-level retrieval across all content types
✅Content ownership model with designated owners, review cadences, and staleness indicators per document
✅Modular runbook templates eliminating duplicate authoring — write once, reuse across teams
✅Sandbox-validated Okta/SAML guides — 98% production deployment QA pass rate; zero credential incidents post-deployment
✅Average retrieval time: 3 minutes — 75% improvement
Information Architecture — Content Strategy Taxonomy
Unified content hierarchy designed and implemented across all teams · Content type tags drive semantic search classification and Knowledge Governance ownership rules
📚 DoorDash IT Knowledge Base
├── 🚨 Incident ManagementTask
├── RunbooksTask
├── P0 Incident Response Runbook [template]
├── Service Degradation Runbook [template]
└── Rollback Procedures [template]
├── Escalation PathsReference
└── On-Call Rotation & Contact Matrix
├── Jira Incident Ticket ConventionsReference
└── Post-Incident ReviewsConcept
├── 🔐 Security & AuthenticationTask
├── SSO / Okta Integration GuidesTaskNew
├── Okta SAML Integration — Sandbox Setup Guide
├── SAML Assertion Troubleshooting Reference
└── Credential Rotation Procedures
├── Authentication FlowsConcept
└── Security Incident RunbooksTask
├── 🏗️ Platform & ArchitectureConcept
├── Service CatalogReference
└── [Service Name] — Owner, Dependencies, Endpoints
├── System Design DocumentsConcept
├── API & SDK DocumentationReferenceNew
├── Internal API Reference — Endpoints, Auth, Rate Limits
└── SDK Integration Guides — Setup, Usage, Versioning
└── CI/CD Pipeline DocumentationTask
└── Jenkins Pipeline Runbooks & Deployment Procedures
├── 🚀 Engineering EnablementGolden Path
├── Golden Path — New Engineer OnboardingGolden PathNew
├── Day 1: Environment Setup
├── Week 1: First Deployment
└── Month 1: Service Ownership
├── Tooling Setup GuidesTask
└── Development StandardsReference
└── ⚙️ IT OperationsTask
├── Change ManagementTask
├── Deployment ProceduresTask
└── Monitoring & Alerting PlaybooksTask
✍️ Product UX Writing Decision — Golden Path Architecture
The Golden Path onboarding structure applies product UX writing principles to developer experience: progressive disclosure (Day 1 → Week 1 → Month 1), task sequencing aligned to cognitive load, and a single canonical path that eliminates the 3-competing-guides problem surfaced in the audit. Each stage is scoped to one decision the engineer needs to make — not a list of everything they could possibly configure. This is the same information architecture principle that drives in-product onboarding design, applied to internal Developer Experience (DX) documentation.
ML-informed Semantic Search — Parameter Specification Architecture
Designed the retrieval specification for an ML-enhanced search layer across the unified knowledge base · Engineered and validated with Python test pipeline · Click any parameter to expand
AI-AUGMENTED WORKFLOW
🔎 ML-informed Retrieval — Search Parameter Specification
Standard keyword search was insufficient because engineers under incident pressure search by intent ("how do I restart the auth service") not by document title ("SSO Service Restart Runbook v2.1"). I designed the full specification architecture for an ML-informed semantic search layer — Confluence's semantic ranking uses ML-based relevance scoring — and engineered the parameter logic to shape how that model routes queries to content. I authored the search parameter specification document as the engineering brief: defining each parameter's values, weighting rationale, and expected retrieval behavior. A platform engineer then implemented the parameters against the Confluence ML search layer, using my spec as the implementation reference. I validated retrieval accuracy against a test query library drawn from real incident reports, iterating on the parameter logic until retrieval met the 3-minute target.
valuesrunbookconceptreferencegolden-path
weight1.8 — boosted above title match to prevent keyword collisions across content types
rationaleEngineers searching "restart" during an incident should surface runbooks, not concept articles about restart patterns. Type classification enforces this intent routing before keyword scoring.
trigger_termsincident outage down failing on-call alert P0 P1 rollback
behaviorWhen urgency_context = true, results filter to content_type: runbook and re-rank by last_validated date descending — ensuring on-call engineers always see the most recently verified procedure first.
impactEliminated the primary failure mode: engineers reading stale runbooks during live incidents because newer versions ranked lower by page-view score.
valuesincident-response security-eng platform-eng sre devops it-ops eng-enablement qa-eng
weightTeam-scoped results receive +0.4 relevance boost. Cross-team results remain visible at lower rank — preserving discoverability across boundaries.
rationaleHard-filtering by team recreated the silo problem. Soft boosting gives engineers their team's content first while surfacing platform-wide resources when local results are insufficient.
logicIf today - last_reviewed_date exceeds review_cadence threshold, apply relevance penalty of -0.6 and surface a staleness badge in search results UI.
cadence_defaultsrunbook → 90 days · concept → 180 days · reference → 365 days · golden-path → 60 days
rationaleThe audit found that stale documents ranked highly by page-view score — engineers navigating to them during incidents compounded the staleness problem. This parameter inverts that reward structure.
examples SSO → single sign-on, Okta, SAML, auth, authentication
runbook → playbook, procedure, SOP, how-to
restart → reboot, cycle, recover, remediate
incident → outage, degradation, alert, P0, P1
rationaleAudit revealed each team used different terminology for identical concepts. Without synonym normalization, a platform-eng engineer searching "playbook" would miss incident-response content titled "runbook" — and vice versa.
Content Model — Four Type Classification System
Drives both search parameter routing and governance ownership rules
🧭
Concept
Explains what something is or how it works. No action steps.
review: 180d · owner: arch team
📋
Runbook / Task
Step-by-step procedure for a specific operational task or incident response.
review: 90d · owner: responsible team
📊
Reference
Lookup tables, service catalogs, contact matrices, API specs.
review: 365d · owner: named DRI
🚀
Golden Path
Canonical onboarding path for a new engineer or service. Curated and endorsed.
review: 60d · owner: eng-enablement
Measured Outcomes
Post-implementation metrics vs. pre-consolidation baseline
75%
Faster Incident Retrieval
12 min → 3 min average retrieval under incident conditions
20+
Engineering Hours Reclaimed / Week
Equivalent to 0.5 FTE of recovered capacity · Reduced cognitive load and lowered barrier to entry for new engineers
$180K+
Annual Savings
Via modular runbook reuse and reduced documentation overhead
98%
QA Pass Rate
Production deployment pass rate for sandbox-tested Okta/SAML integration guides
DX+
Developer Experience
Reduced time-to-productivity for new engineers · Single Golden Path eliminated 3 competing onboarding guides · Lower DX friction across 6 teams
📦 Consolidation Summary
📁25+ repositories consolidated into 1 centralized platform
🔐Zero credential-related incidents post auth-guide deployment
🌲4-type content model adopted across all 6 IT engineering teams
🔍5 semantic search parameters implemented with engineering team
📋Modular runbook templates deployed; reuse across 6+ teams
💰 Savings Breakdown — $180K+ Annual
Recovered engineering time — 20+ hrs/week across 6 teams via faster retrieval, reduced onboarding friction, and elimination of duplicate documentation overhead $94K+
Reduced duplicate authoring via modular runbook templates (reuse across 6+ teams) $48,000
Prevented credential incidents (avg $19K per incident) $38,000+
Modular Runbook Template — Standard Structure
Required fields enforce consistency across teams · Optional fields support team-specific extensions without breaking the shared model
📋 Incident Runbook Template — v2.0 Modular content_type: runbook · review: 90d
● Module A — Identity & Context
Document Title[Service Name] — [Action] RunbookRequired
Owner / DRI@username · team-nameRequired
Last ValidatedYYYY-MM-DD · validated in: staging | productionRequired
Severity LevelP0 | P1 | P2 | P3Required
Linked ServicesService catalog IDs — comma-separatedRequired
● Module B — Trigger & Symptoms
Alert NameExact alert string from monitoring systemRequired
SymptomsObservable behavior that triggers this runbookRequired
False Positive RateKnown conditions that trigger alert without actual incidentOptional
● Module C — Diagnostic Steps
Step 1Confirm alert validity: [specific check + expected output]Required
Step 2Identify scope: [dashboard link + what to look for]Required
Step NDecision tree: [condition] → [action] | [condition] → [escalate]Required
● Module D — Resolution & Recovery
RemediationExact commands with expected outputs — copy-paste readyRequired
ValidationHow to confirm the issue is resolved before closingRequired
RollbackIf remediation worsens the incident: [rollback steps]Required
Post-IncidentLink to PIR template · auto-populated fieldsOptional
Knowledge Governance — Content Ownership & Review Model
Every document in the unified platform carries a designated owner, a review cadence, and a quality gate threshold — enforcing the Knowledge Governance standard across all 6 engineering teams
Content Type Primary Owner Review Cadence Staleness Trigger Quality Gate
Runbook Responsible Team Lead Every 90 days Any service dependency change; any related incident; 90-day elapsed Sandbox-validated
Concept Architecture Team / TL Every 180 days Major architecture change; product pivot; 180-day elapsed SME reviewed
Reference Named DRI (per doc) Every 365 days Team restructure; tool change; contact update needed Accuracy check
Golden Path Eng. Enablement Lead Every 60 days Any new-hire feedback; tooling change; onboarding friction report Onboarding-tested
Transformation Workflow — 5-Phase Implementation
From audit through governance rollout · Tools at each phase
1
Repository Discovery & Inventory
I catalogued all 25+ documentation sources by team, tool, page count, last-modified date, and ownership status. I classified each source by content type and flagged documents with no owner or review date for immediate remediation.
Confluence GitHub Jira Notion Google Drive Audit spreadsheet
2
Stakeholder Interviews & Pain Point Mapping
I designed and ran structured discovery sessions with 6 engineering teams, using a consistent interview framework: current workflow, point of friction, last time they couldn't find something, and what they needed most. I synthesized findings into the 4-theme pain point matrix.
Interview framework Friction-impact matrix Affinity mapping
3
Taxonomy Design & Content Model
I designed the 4-type content classification system and unified IA taxonomy, establishing naming conventions, folder hierarchy, and content model metadata schema (owner, review date, content type, team scope, linked services). I presented the architecture to team leads for review and sign-off before implementation.
IA taxonomy Content model doc Naming conventions Confluence Jira
4
ML-informed Semantic Search — Parameter Specification & Validation
I designed the full specification for 5 semantic search parameters (content_type, urgency_context, team_scope, staleness_penalty, synonym_map) and authored the parameter specification document that served as the engineering brief — translating content strategy intent into implementation logic. A platform engineer implemented the parameters against Confluence's ML-enhanced search layer using my spec as the reference. I then validated retrieval accuracy against a test query library drawn from real incident reports, iterating on the parameter logic until the 3-minute retrieval target was achieved.
Search parameter spec Confluence CQL / ML ranking Test query library Jira SIT validation
5
Governance Model Rollout & Template Deployment
I deployed modular runbook templates and the Knowledge Governance ownership model across all 8 teams, and trained team leads on review cadences, quality gates, and staleness workflows. I authored Okta/SAML integration guides through a full sandbox-to-production cycle, achieving a 98% QA pass rate on first deployment. The combined impact of recovered engineering time, modular template reuse, and prevented credential incidents established $180K+ in annual savings.
Runbook templates Okta SAML Jenkins Jira QA testing (SIT/UAT) Governance docs
Quality Gates — Documentation Standards
"Value over volume" — every document published must clear these gates before entering the knowledge base
✓
Content Type Classification
Every document must be classified as Concept, Runbook, Reference, or Golden Path before publication. Classification determines review cadence, ownership rules, and search parameter routing.
Gate: content_type metadata field must be populated
✓
Designated Owner (DRI)
Every document must have a named owner — a specific person or team alias, not a group. Anonymous documents cannot be published; ownerless documents in migration are flagged for immediate assignment.
Gate: owner field ≠ null · team alias accepted, generic titles rejected
⬡
Runbooks: Sandbox Validation Required
All runbooks and integration guides must be executed end-to-end in a non-production environment before the validated date can be set. This gate was specifically designed after audit findings revealed that untested Okta/SAML guides had caused production credential incidents.
Outcome: 98% production deployment QA pass rate · Zero credential incidents post-deployment
⬡
Modular Template Compliance
Runbooks must use the standard modular template (Modules A–D). Required fields must be populated. Optional fields may be omitted but cannot be removed from the template structure, preserving compatibility with the search parameter schema.
Impact: Eliminated duplicate authoring · Enabled content reuse across 6+ teams · $48K annual savings
✓
Naming Convention Compliance
Document titles must follow the [Team/Service] — [Action/Topic] — [Type] naming pattern. Enforces consistency in search result display and prevents the synonym collision problem identified during audit.
Gate: Title parser flags non-compliant names before publish · Auto-suggestion provided