Communicating to Technical and Executive Audiences
Architecture communication is explaining a system so each audience can make the decision that belongs to them: executives decide whether to fund and accept the risk, engineers and security reviewers decide whether the design is sound, end users decide whether to trust it. This note covers answer-first structure, telling the same system two ways, explaining uncertainty and eval results to non-specialists, whiteboarding a design, demos that build justified trust, and handling the objections every AI project meets.
- Structure a message answer-first (conclusion, then supporting points, then evidence)
- Explain the same AI system to an executive and to an engineering reviewer, choosing what each needs
- Present eval results and uncertainty in plain language without overstating or hiding them
- Run a whiteboard design discussion and a demo that shows limits as well as strengths
- Answer the common objections to AI projects with evidence rather than reassurance
- Architecture Docs & ADRs
- Building Your Own Evals - confidence intervals
Lead With the Answer
Barbara Minto's Pyramid Principle - and the military's "bottom line up front" - say the same thing: start with the conclusion, then the few arguments that support it, then the evidence behind each. Busy readers stop early; give them the part they need first.
flowchart TD
A["🎯 Answer<br/>Fund a 12-week pilot of the reply assistant for the UK team"]
A --> B1["💸 It pays back<br/>3-7 months across our scenarios"]
A --> B2["✅ It works on our tickets<br/>91% of drafts usable in testing"]
A --> B3["🛡️ The risk is contained<br/>agents review every draft"]
B1 --> E1["Cost model,<br/>sensitivity table"]
B2 --> E2["300-ticket eval,<br/>failure examples"]
B3 --> E3["Threat model,<br/>rollback plan"]
style A fill:#e8e0d4,stroke:#c8b89a
style B1 fill:#d8dfe8,stroke:#b0bac8
style B2 fill:#d8dfe8,stroke:#b0bac8
style B3 fill:#d8dfe8,stroke:#b0bac8
| Bottom-up (hard to follow) | Answer-first |
|---|---|
| "We evaluated three models on 300 tickets, built a retrieval index, measured edit times, modelled costs under four scenarios, and therefore we recommend a pilot." | "We recommend a 12-week pilot. It pays back in 3-7 months, drafts were usable 91% of the time on our own tickets, and agents review everything before it is sent. Details follow." |
One System, Two Audiences
| Executive / sponsor | Engineering / security reviewer | |
|---|---|---|
| Wants to decide | Fund it? Accept the risk? | Is the design sound, secure and operable? |
| Cares about | Outcome, cost, timeline, risk, what happens if it fails | Data flows, failure modes, dependencies, evals, operations |
| Unit of evidence | Business metric moved, payback, comparable cases | Eval results with CIs, threat model, load tests, ADRs |
| Time you get | 5-10 minutes, often interrupted | An hour, with questions |
| Avoid | Jargon, architecture diagrams, model names | Hand-waving, marketing language, untested claims |
Here is the same reply assistant explained both ways. Use the audience toggle at the top of the page to switch views.
For the sponsor: The assistant writes a first draft of each customer reply using our own policy documents and the customer's history. An agent reads it, edits it and sends it - nothing goes out without a person approving it. In testing on 300 real tickets, about nine in ten drafts needed only small edits, cutting the time per reply from about 9 minutes to about 3. At the adoption we expect, that frees roughly 2,000 agent hours a month and pays back the build in three to seven months. The main risk is a draft with wrong policy information; agents are the check on that, and we measure how often they catch problems. If the pilot doesn't hit its targets after 12 weeks, we stop.
For the reviewer: A ticketing plug-in calls an assistant API, which runs hybrid retrieval (BM25 plus dense vectors, reranked) over a nightly-built policy index and the ticket history, then calls a hosted model through our gateway with a fallback provider. Drafts carry citations to policy sections. Customer email content is treated as untrusted data: the model has no write tools, so injection can at worst produce a bad draft that an agent reviews. Evaluation: 300 held-out tickets, 91% "send with minor edits" (95% CI 88-94%), groundedness checked by a judge validated against 100 human labels; the same judge runs online on 5% of traffic. Traces follow OpenTelemetry GenAI conventions; SLOs are TTFT p95 under 1.5 s and 99.5% availability with a retrieval-only degraded mode. Decisions and their revisit triggers are in ADRs 0001-0006.
Explaining Uncertainty and Eval Results
Executives make decisions under uncertainty every day; what they need is the uncertainty stated plainly, not hidden or drowned in statistics.
| Instead of | Say |
|---|---|
| "91.3% accuracy" | "About 9 in 10 drafts needed only small edits on 300 of our own tickets. The true rate is very likely between 88% and 94%." |
| "Hallucination rate 2.1%" | "About 1 in 50 drafts contained a policy detail that wasn't in our documents. Agents caught all of them in testing; we'll keep measuring that in the pilot." |
| "It's state of the art" | "It beat the other two options we tested on our tickets, and here's how it compares to how our agents do today." |
| "It never makes mistakes" | (Never say this.) |
Practices that keep the message honest:
- Use natural frequencies ("1 in 50") rather than percentages for rare events - research by Gigerenzer and Hoffrage shows people reason about them more accurately.
- Compare with the human baseline. "Agents mis-quote policy in about 3% of replies today" turns an abstract error rate into a decision.
- Show real failures. Two or three actual bad outputs, and how each would be caught, build more trust than a perfect score.
- Say what you haven't tested - other languages, rare ticket types, peak load.
Whiteboarding a Design
Whether in an interview or a customer workshop, the same sequence works:
- Clarify - restate the goal, the users, the scale and the constraints; ask what success looks like. Write the numbers on the board.
- Draw the context - users, the system as one box, the systems around it.
- Walk the main flow - one request, end to end, adding containers as you go.
- Find the hard part - usually retrieval quality, an integration, a risk control or a latency budget - and go deep there.
- State the trade-offs - "We could self-host for cost, but at this volume the hosted API is cheaper once you count on-call; I'd revisit at about 10x the volume."
- Close with how you'd know it works - evals, SLOs, the pilot metric.
Think out loud, check in with the room ("does this match how your team works?"), and treat a correction as information, not a defeat.
Demos That Build Trust
A demo should leave the audience with an accurate picture, not just a good feeling - an over-sold demo turns into a failed pilot when reality arrives.
- Use their data - real, anonymized examples from discovery, including a messy one.
- Show the human control - the review step, the approval, the citation the user can click.
- Show a failure on purpose - an out-of-scope question, a missing document - and what the system does about it.
- Be honest about selection - if examples are hand-picked, say so, and show the eval numbers that describe the typical case.
- Have a backup - a recording, in case the live service or network fails.
Handling Objections
| Objection | What's usually behind it | Answer with |
|---|---|---|
| "It makes things up." | Fear of a public, costly error | Measured error rate vs the human baseline, how errors are caught (review, citations, guards), what the system does when unsure |
| "Our data will leak or be used for training." | Security, legal and contractual exposure | Data-flow diagram, provider terms (retention, training use), region, access controls, the threat model |
| "It's too expensive." | Unclear value or unpredictable cost | Unit cost model, payback range, spend caps and alerts, the cheapest option that met quality |
| "It will replace my team." | Job security, loss of control | The workflow role chosen (inform, assist) and why; what the team will do with the time; their involvement in the design and the pilot |
| "We'll be locked in." | Vendor dependence | Model gateway, eval set that makes switching testable, ADR with the revisit trigger |
| "How will we know it still works next month?" | Silent degradation | Online evals, SLOs, version pinning, alerting, incident process |
| "Who is accountable when it's wrong?" | Legal and organizational responsibility | The human decision point, audit logs, the process owner named in discovery |
Answer with evidence and design, not reassurance. An objection you can't answer is a risk to add to the register - and saying "we don't know yet, here's how we'll find out" is a credible answer.
Check Yourself
- You have 5 minutes with a CFO to get approval for a pilot. What do you say first?
- Which statement presents an eval result best to a non-technical sponsor?
- Why include a deliberate failure in a demo?
- A team lead says 'This will replace my team.' What is a good response?
Exercises
Rewrite this status update for a steering committee, answer-first, in under 80 words:
"This sprint we finished the ingestion pipeline for the second document source, which took longer than expected because of permission mapping issues. We also re-ran the eval set and accuracy went from 84% to 89% after the reranker change. Latency went up by about 300 ms. We think we'll need another two weeks before the pilot can start, and we need a decision on whether the legal team's documents can be included."
Solution
Pilot start moves two weeks, to 14 November, and we need one decision: can legal's documents be included? Progress: accuracy rose from 84% to 89% with a new reranker (latency +300 ms, still inside target). The delay came from mapping document permissions for the second source, now done. Decision needed by 25 October from the General Counsel's office.
You are presenting a contract-review assistant to a law firm's partners. Prepare a one-sentence answer, with the evidence you would bring, for: (1) "It will miss a clause and we'll be liable." (2) "Client documents can't leave our environment."
Solution
- "It flags clauses for a lawyer to review rather than approving anything, and in our test on 120 past contracts it found 96% of the clauses your associates flagged, plus 9 they missed - here are the misses on both sides." Evidence: the eval with per-clause-type recall, examples of misses, the review workflow.
- "Processing runs in your own cloud tenant with no provider retention or training on your data, and here is the data-flow diagram your IT team reviewed." Evidence: data-flow diagram, provider terms, the security review sign-off, access logs.
Study Notes
Must-know:
- Answer first (Pyramid Principle / bottom line up front): conclusion, supporting points, evidence
- Executives decide on value, cost and risk; engineers and security decide on soundness - same facts, different selection and order
- Explain uncertainty in natural frequencies, with the sample and the range; compare with the human baseline; show real failures
- Whiteboard sequence: clarify with numbers, context, main flow, hard part, trade-offs, how you'd know it works
- Demos: their data, the human control, a deliberate failure, honest selection, a backup
- Answer objections with evidence and design; an unanswerable objection is a risk for the register
References
- Minto, The Pyramid Principle (1987; 3rd ed. 2009)
- Gigerenzer and Hoffrage, How to Improve Bayesian Reasoning Without Instruction: Frequency Formats (Psychological Review, 1995)
- Amershi et al., Guidelines for Human-AI Interaction (CHI 2019)
- Google PAIR, People + AI Guidebook - Explainability + Trust (2021)
- Bezos, 2017 Letter to Shareholders (Amazon, 2018) - narrative memos instead of slides
Last reviewed: 2026-10