All work

Case study 01 GSA / IT Vendor Management Office

Conversational AI for Federal Navigation

A chatbot that could only answer from approved sources, and the trust problem that created.

Client
GSA / IT Vendor Management Office
Role
Research, information architecture, design standards, conversational AI
Timeline
Three months, 2023
Team
One designer, one developer, program stakeholders
68% reduction in task friction
77 → 25 PURE score across three rounds
Beta built, not deployed

Context

Everything was published. Nothing was findable.

The IT Vendor Management Office helps federal buyers make sense of IT acquisition. Contracting officers, program managers, and small-business specialists arrive with real decisions in front of them: which vehicle covers a given service, whether a set-aside applies, what a pricing escalation clause permits.

The information was all there. 221 resources across seven topic areas. Technically, everything a buyer needed was on the site.

Content inventory

221 resources, seven topic areas

Counts exceed the total because resources carry more than one tag.

  • Acquisition best practices 52 resources tagged Acquisition best practices
  • Market intelligence 46 resources tagged Market intelligence
  • Contract solutions 38 resources tagged Contract solutions
  • Policy 34 resources tagged Policy
  • Technology 29 resources tagged Technology
  • Small business 28 resources tagged Small business
  • ITVMO general 24 resources tagged ITVMO general

The problem

A published document is not an answerable question.

The site was organized around how the agency filed its content, not around what a buyer was deciding. Finding an answer meant knowing which of seven categories it lived under, and buyers do not think in those categories.

In federal acquisition, a confidently wrong answer is worse than no answer.

A buyer who acts on invented guidance about procurement rules has a compliance problem, not a usability problem. That ruled out the obvious approach before it started.

Constraints

What shaped every decision that follows.

Three months
Against an existing live site.
One developer
Every decision had to be buildable by a single person inside the timeline.
Non-technical maintainers
Business analysts and project managers needed to update content afterwards without a developer, so the CMS was a design constraint rather than an implementation detail.
Section 508 and WCAG
Applied to a conversational interface where the patterns are less settled than they are for pages.
No invented answers
The system could not reach past what it had been given.

Discovery and research

Two inputs, pulling in the same direction.

Google Analytics gave the behavioral picture: where buyers entered, what they searched for, where sessions ended without a result. It showed what was happening but not why. So I built the qualitative side alongside it: personas, journey maps, and a full architecture audit of what existed against what buyers were trying to do.

Buyer journey

The site got them close. Close was not the answer.

  1. Step 1

    Arrives with a specific question

    A vehicle, a clause, a set-aside. One decision, already in mind.

  2. Step 2

    Task-based navigation gets close

    The redesigned architecture lands them on the right topic in a click or two.

  3. Step 3, stall point

    The resource is close, not exact

    Adjacent guidance. A different vehicle, year, or clause.

  4. Step 4

    Re-reads, hedges, or gives up

    No way to confirm “close enough” without finding a person to ask.

The same two stalls repeated across nearly every mapped path: content that was topically close but procedurally different, with no way to confirm it before acting.

Read together, the two agreed. Analytics showed sessions ending on category pages. The journeys showed why: a buyer arrives with a decision, and the site asked them to already know which of seven filing categories the answer lived under.

To measure it I used the PURE method, a structured expert evaluation that scores each task by the difficulty a qualified user would face. It suits situations where direct user access is limited, which this was: federal buyers are hard to recruit and three months did not allow for a full study. I scored seven representative buyer tasks three times. A baseline in June, a mid-build check in August, and a final round against the redesigned site in November.

PURE gave me expert judgment about how hard a task should feel. Analytics gave me what people were actually doing. The two disagreed in useful places: content that scored acceptably still showed heavy drop-off, which usually meant buyers were arriving from search and landing mid-hierarchy with no orientation.

PURE evaluation

Seven buyer tasks, three rounds of friction scores

Accessible matrix of seven representative buyer tasks evaluated in June, August, and November 2023. Each task has five colored segments representing the friction of its steps. Lower totals indicate less friction; the round totals fall from 77 to 53 to 25.

Low friction Moderate friction High friction
PURE evaluation by buyer task and round
Buyer task June June 29, 2023 77 total friction score August August 15, 2023 53 total friction score November November 7, 2023 25 total friction score
Find a contract vehicle
June: 10 points. high friction, high friction, high friction, moderate friction, moderate friction.
August: 7 points. moderate friction, moderate friction, low friction, low friction, low friction.
November: 3 points. low friction, low friction, low friction, low friction, low friction.
Find technical standards
June: 14 points. high friction, high friction, high friction, moderate friction, moderate friction.
August: 5 points. moderate friction, low friction, low friction, low friction, low friction.
November: 3 points. low friction, low friction, low friction, low friction, low friction.
Learn about the Polaris contract
June: 15 points. high friction, high friction, high friction, moderate friction, moderate friction.
August: 6 points. moderate friction, low friction, low friction, low friction, low friction.
November: 3 points. low friction, low friction, low friction, low friction, low friction.
Discover small-business opportunities
June: 18 points. high friction, high friction, high friction, high friction, high friction.
August: 9 points. high friction, moderate friction, moderate friction, low friction, low friction.
November: 8 points. moderate friction, moderate friction, low friction, moderate friction, low friction.
Understand set-aside requirements
June: 8 points. high friction, high friction, moderate friction, moderate friction, moderate friction.
August: 8 points. moderate friction, moderate friction, low friction, low friction, low friction.
November: 3 points. low friction, low friction, low friction, low friction, low friction.
Compare contract solutions
June: 7 points. high friction, high friction, high friction, moderate friction, moderate friction.
August: 7 points. moderate friction, moderate friction, low friction, low friction, low friction.
November: 2 points. low friction, low friction, low friction, low friction, low friction.
Find a pricing escalation clause
June: 5 points. high friction, high friction, moderate friction, moderate friction, low friction.
August: 11 points. moderate friction, low friction, low friction, low friction, low friction.
November: 3 points. low friction, low friction, low friction, low friction, low friction.
Each column is a round, each row a task. Segments show the difficulty of each step: green is low friction, amber is moderate, and red is high.

How the work was made

Foundations, then fidelity, then sign-off.

  1. Personas and journeys

    I mapped how contracting officers, program managers, and small-business specialists move from a question to a decision. Those journeys became the test for every structural decision: if a proposed architecture did not shorten one of them, it did not go in.

  2. Site architecture

    The architecture came out of the journeys rather than the existing content taxonomy. That is the whole move in this project. The old structure described how ITVMO files things; the new one describes what a buyer is trying to decide.

  3. Low fidelity first

    I stayed in wireframes for as long as the questions were structural. A stakeholder looking at a gray wireframe argues about whether the structure is right; a stakeholder looking at a polished screen argues about the shade of blue.

  4. Prototype as specification

    Once structure settled I moved to high-fidelity interactive prototypes in Axure, so the developer could see behavior rather than infer it from a static comp. On a compressed timeline the round trip is the expensive part.

  5. Stakeholder review as the gate

    Nothing was built before program stakeholders approved it. With one developer and three months, a late objection would have been unrecoverable, so clearing each stage before moving on was what made the timeline hold.

What I tried first

The simpler fix, because it was clearly broken.

I re-architected navigation and wayfinding around buyer tasks instead of agency filing categories, built a content-pattern library so every resource type had a consistent shape, and wrote plain-language guidance for the people producing content.

Information architecture

Filed by category, or organized by task

Before

The buyer must know which of seven categories the answer lives under.

  1. Home
  2. Resources
  3. Topic area
  4. Filter set
  5. Resource

Five steps, and step three requires a guess.

After

The buyer states the decision they are making and is routed to the answer.

  1. Home
  2. Task path (key step)
  3. Answer

Three steps, and no category knowledge required.

We built on Netlify with a git-based CMS so business analysts and project managers could publish events and add resources themselves, with no ticket and no developer in the queue. That is the reason the content patterns mattered: a CMS that lets anyone publish anything recreates the original problem inside a year. The patterns constrained publishing into shapes the architecture already accounted for, and I ran the training that got the team using them.

Standards

Fixing the site was not enough on its own.

ITVMO publishes fact sheets, slide decks, and reports as well as web pages, and all of it was drifting. A buyer who found a consistent answer on the site and then downloaded an inconsistent PDF was back where they started.

So I authored the ITVMO Document Style Guidelines: a logo system with spacing rules and misuse examples, 508 color-contrast guidance with pass and fail cases, a three-typeface hierarchy, primary and secondary palettes with hex values and shade ramps, and full specifications for fact sheets and slide decks down to margins, line spacing, and paper weight.

Twenty-four pages. The 508 section does not just state the requirement. It shows six logo-on-background combinations marked pass or fail and links a contrast checker, so anyone producing a document can test their own work.

The constraint

It solved one problem and created another.

For certain task types the site was faster but still indirect. A buyer with a specific question about a specific vehicle already knew what they wanted to know. Making them traverse a well-organized hierarchy was still making them translate their question into someone else’s structure. That pointed toward a conversational layer, and in federal acquisition a chatbot is a liability as much as a feature.

The system was restricted to retrieve only from an approved document set. It could not reach past what it had been given, and could not generate an answer that was not grounded in a source. That solved accuracy completely. It also created the design problem that became the actual work.

A source-constrained system says it does not know far more often than a general-purpose one. How it handles that moment determines whether anyone uses it twice.

Four states, and only one is the clean answer.

  1. The grounded answer

    Every response carries its source, visible before the user acts. A contracting officer putting their name on a decision needs to check the guidance, and a citation they have to hunt for is a citation they will not use.

  2. The partial answer

    Sources cover part of the question. The system says which part, rather than filling the gap or refusing entirely.

  3. The gap

    No approved source. This is the state most conversational design skips, and it is where trust is won or lost. The system says so plainly and routes to a person. No apology loop, no rephrasing suggestions that would not have worked, no confident guess.

  4. The recovery

    The first answer missed. A route forward that is not starting over.

Underneath all four: latency. Retrieval takes time, silence reads as broken, and a spinner reads as stalled. The system had to show what it was doing.

An empty field is a hard interface.

The assistant opens with four routes: find events, explore resources, learn about ITVMO, watch past events. Alongside them sit suggested questions drawn from what buyers actually ask. That reduced the blank-input problem and set expectations about scope before anyone hit a wall.

The framework

The trust patterns generalized past this project.

I co-authored a responsible-use framework with ethics, compliance, and risk, covering transparency, latency feedback, error handling, and user-trust design across conversational interfaces, automated guidance, and AI-assisted navigation. Built so the next team would not relitigate the same questions.

Responsible-use framework

Four principles, one standard everywhere.

Co-authored with ethics, compliance, and risk. Applied identically across all three surfaces.

  1. Transparency

    Where an answer came from, visible before a user acts on it.

  2. Latency feedback

    What the system shows while it is working, so silence does not read as broken.

  3. Error handling

    What happens when a source does not cover the question: no invented answers, no dead ends.

  4. User-trust design

    How recovery works after a wrong or partial answer, so one miss does not end the session.

Surfaces covered
  • Conversational interfaces
  • Automated guidance
  • AI-assisted navigation
Built so the next team would not relitigate the same questions.

What came of it

Friction fell 68%, and the assistant never shipped.

Ask Vemo, the ITVMO assistant. Beta build, 2023. Note the persistent line at the foot of the panel: answers based on official ITVMO content, with sources one click away.

Measured outcome

Task friction fell 68% across three evaluations

PURE method. Lower is better. Each round scored the same seven buyer tasks.

  1. 77 PURE friction score June 29, 2023
  2. 53 PURE friction score August 15, 2023
  3. 25 PURE friction score November 7, 2023
Biggest movers: finding technical standards fell 14 to 3, and learning about the Polaris contract fell 15 to 3.

Every one of the seven tasks improved. The largest movers were finding technical standards for a new acquisition, which fell from 14 to 3, and learning about the newly BIC-designated Polaris contract, which fell from 15 to 3.

Shipped

  • Information architecture organized around buyer tasks
  • Content-pattern library and plain-language guidance
  • Git-based CMS the program team publishes to without a developer
  • ITVMO Document Style Guidelines, twenty-four pages

Did not ship

  • The assistant was completed to beta and never deployed
  • The budget for it did not survive, a funding decision rather than a design one

The retrieval-constrained approach matched what the industry now calls RAG. I was designing for it before the vocabulary settled. The architecture matched the pattern, not that I built a RAG system to current standards.

What I would do differently

Three things.

  • I would have tested the uncertainty states specifically

    I designed them from principle and from what the retrieval constraint implied, and I believe they were right. But I never watched a contracting officer hit “I don’t have an approved source” and decide whether to trust the system afterward. That is the moment the whole design turns on, and I have reasoning for it rather than evidence.

  • I would have separated the two projects earlier

    The IA work and the conversational layer were treated as one effort because they shared a goal, and the assistant’s funding got tied to a scope it did not need. Split, the assistant might have found a smaller budget line of its own.

  • One task never got where it should have

    Discovering small business opportunities started worst at 18 and ended highest at 8. Every other task finished at five or below. I ran out of runway before I solved that one, and it is the task with the most public-policy weight attached to it.