Avery

An agentic AI assistant combining knowledge agents, insight reports, and automated workflows.

*Some content has been redacted due to PII and confidentiality.

Avery — assistant home surface

01 — Context

Context

Role
Design Team Lead
Timeline
~9 months
Tools
Figma, FigJam, Figma Make, Copilot, Dscout
Team
7 designers, 35 engineers, 6 PMs

02 — Discovery

Discovery

I started where the pain was loudest: the long tail of repetitive questions users kept routing to support. We pulled six months of transcripts and tagged them by intent, which sorted the mess into eight capability areas an agent could plausibly own.

Then I ran 12 one-on-one sessions across power users, casual users, and internal operators. The same two asks came up in nearly every conversation: people wanted the assistant to remember context across turns, and they wanted to see where an answer came from. Nobody used the word "trust," but that was the thing they were describing.

03 — Problem & Alignment

Problem & Alignment

The chat experiences we were competing with treated every request as a one-shot prompt — no memory, no sources, and no honest signal for the difference between "I'm guessing" and "I checked the source of truth." Engagement looked healthy; trust didn't. The follow-ups it generated quietly cancelled out the time it was supposed to save.

With the AI/ML team and product leadership, we landed on a multi-agent model: a routing layer in front of eight specialized capabilities (knowledge lookup, reporting, commission insights, digital service tracker, plan recommender, universal census reader, renewals, and quoting), each with its own evaluation harness and confidence signal.

The constraint I held onto: every response had to make the agent's identity, sources, and confidence legible — without turning the interface into a dashboard.

04 — Strategy & Product Plan

Strategy & Product Plan

The bet: a single conversational surface where it's obvious the assistant is working on your behalf — retrieving, reasoning, deferring — not just producing text.

What "good" meant: each capability had to feel trustworthy enough that people acted on its answers without re-checking, and resolve most requests without a chain of follow-ups. We validated this in moderated sessions rather than chasing a single satisfaction number.

The design system for it:a consistent agent-response anatomy — scoped action, source chips, confidence band — reused across all eight capabilities; a shared "what I can do" overview; and a transparent handoff for the moments the agent should step back and let a human take over.

How we kept it honest: weekly prompt-and-UI critiques with the ML team, fortnightly moderated tests with 4–6 participants, and post-launch dashboards tracking resolution per capability.

05 — Takeaways & Outcomes

Takeaways & Outcomes

We shipped the assistant across all eight capabilities. In moderated testing it cleared the trust bar we'd set, time-to-answer on the long-tail intents dropped noticeably, and the source-attribution pattern got picked up by adjacent surfaces looking to solve the same credibility problem.

What stuck with me: people didn't need the assistant to be right more often — they needed to see why it was confident. The trust came from showing the work, not from a better model.

Selected screens

Homepage
Homepage
Expanded
Expanded
Collapsed
Collapsed

Let's talk.

A little about me: