Designing AI experiences for deep analysis and traceability.
While leading the design of PANW’s Sales Workbench, I shaped the user experience for two critical AI features.
Anchored Follow-Ups help users converge on findings while avoiding the drift of long AI-generated narratives.
Trace reimagines “sources” for internal tools while reducing the effort required for verification.
Context:
Sales Workbench consolidated a wide range of tools into one platform where account teams could manage opportunities, accounts, telemetry, and other sales activities. The platform supported oversight of roughly $3B in annual pipeline activity while reducing movement between fragmented tools.
The consolidation also created an opportunity to introduce an AI assistant as an initial step toward AI-assisted workflows.
Problem:
As AI entered enterprise platforms, its narrative exchanges conflicted with the structured, tactile interactions familiar to users.
- An AI assistant was part of the planned experience, but users remained largely indifferent. They considered structured interfaces better suited to the detailed analysis central to their work.
- The effort required to verify an incorrect response continued to outweigh the benefits of AI.
Approach:
Rather than applying AI across the workflow, we identified which tasks could benefit from it and which still required human judgment. Account teams across the Americas, EMEA, and APJC expressed mixed reactions to introducing AI into their workflows.
Takeaway 01
AI is often pitched as delivering “insights,” but information workers need tools that help them examine evidence in depth.
Takeaway 02
Long chat threads increased cognitive load, making it difficult for users to maintain focus and avoid topic drift.
Takeaway 03
Teams often tried to automate work that required human judgment, rather than giving users better tools to uncover hidden challenges.
The focus shifted from generating isolated insights or reports to helping users work through ambiguous questions. AI could synthesize information from multiple sources, but users still needed structured, tactile controls to identify patterns and converge on a finding.
Prototyping and technical hurdles:
With limited engineering capacity, generative UI appeared too costly for V1, so I built a prototype to reduce the uncertainty. Using CopilotKit, it passed the user’s UI context and chat prompt to a Codex shim, which selected an approved React component and returned its structured data.
The prototype gave engineering enough confidence to build a scalable solution on a forked CopilotKit repository.
The verification issue:
Because the data changed constantly, users needed to verify responses quickly. Additional model reasoning added noise.
Transparency was relatively easy; the real challenge was the effort users spent on verification.
Rather than refining the visibility of a model’s thought process, I explored a conclusion-and-derivation response approach. The system responds with an answer and produces a concise, scannable evidence-and-decision trail.
For Trace, I built a taxonomy of process verbs to describe how the machine accessed information, and post-hoc verbs to describe what the model did with the data.
Each source included its last update time because opportunity and account data could change several times a day.
Impact:
Within the first 90 days, one in three participants averaged more than two consecutive Anchored Follow-Ups per session, while one in five averaged more than four. Since each follow-up remained linked to a row or column, they formed connected chains indicating sustained investigation across several steps.
Within the first week, Trace prompted a new feature request: users wanted to capture a response with its trace and share it in Slack, Salesforce notes, or email.
We introduced Snapshot, which reached 27% usage overall, while 22% of anchored follow-up arcs ended with one.
Together with user feedback, these metrics provided our strongest observable signal that roughly one in five investigations reached a useful stopping point, producing a finding worth preserving or carrying into collaboration.
How often users continued their investigation through at least two connected Anchored Follow-Ups.
The average number of Anchored Follow-Ups within each connected investigation.
How often users captured a response and its Trace as an image to preserve or share.
What I learned:
Progress begins somewhere between blind acceptance and complete dismissal.
I was uneasy about adding another chatbot that mimicked thinking and treated citations as decoration. But acknowledging discomfort while taking action led to some of the more unique experiences in the application, with a lasting impact.
I had to shift from asking myself “how do we add a chatbot?” to “what assumptions do we carry about the chatbot experience that might fail our users?”