Datadog Incident Management- Workstreams
Role
Lead Product Designer, UX Research, Prototyper
Platform
WebApp, Slack UI
Shipped
March 2024 (Beta)
CONTEXT
At Datadog (Cloud-based monitoring platform), I’ve been the lead designer on the Incident Management product which is a tool for engineers to investigate and track technical incidents, while collaborating in the product and on Slack to troubleshoot and resolve issues. Team consisted of myself, 1 PM, and 7 ENGs.
PROBLEM
After we GA’ed the Incident Management product, we talked to customers to get feedback and provide support for them and one feature request came up frequently, which was the ability to keep track of multiple workstreams that required different teams in 1 incident, so we ended up calling this feature “Workstreams”.
SOLUTION
This feature became very important to our top leadership and we involved them often throughout the design process to launch Workstreams.
PROCESS
Vision/Goal alignment with stakeholders
Persona Mapping
Sketches
Wireframes and Prototypes 🔁 Validation
Visual Design
Vision / Goal Alignment
In the beginning of a project, I always find it vital to get together with our cross-functional partners (PM and Eng Lead) to define the goals of the project and make sure we’re all in it together on the path forward. Once this is established, I presented the goals to the team as well as to the VP of product during a 1 on 1 sync.
Design Goals
Help ICs (Incident Commanders) track and triage multiple workstreams within a parent incident to resolve rootcauses more quickly and efficiently
Guide Responders to jump into a specific workstream to help resolve issues pertaining to the parent incident
Business Goal
Help customers reduce MTTR (Meant Time To Resolution) of incidents to increase sign ups and revenue of product as a whole
PERSONAS / USER FLOWS
Once project goals have been aligned, I ran multiple research sessions with internal SREs to verify their workflows and dug deeper in their mental state during the different phases.
Below, is the typical user journey of the holistic incident experience while the flow below it illustrates workflows for different personas pertaining to this new feature.
SKETCHES / LO-FI PROTOTYPE / CONCEPT REVIEW
Next, based on my intuitions and experience from talking to internal engs and customers, I sketched out some initial ideas and shared them with my core team - PM and lead Eng. We then agreed on a direction and I quickly mocked up a lo-fi wireframe and prototyped it out and shared it with senior stakeholders.
ROUND 2 WIRES / PROTOTYPES / VALIDATION
Based on the above concept(prototype), senior leadership was excited that this feature may be a market differentiator and gave us the green light to invest more resources to really elevate and make this feature more robust, while building it quickly.
We then proceeded to test hi-fi wireframes and walked through a prototype→(feel free to play with it!) to validate UI concepts and feature preferences with a realistic narrative.
KEY TAKE AWAYS
Majority of the users preferred the card-based collaboration UI, as it gave more screen real estate for each Workstream’s detail with one less click and UI worked better for most incidents
A few users preferred the column-based collaboration UI, as it gave user’s a better bird’s eye view into what’s going on
Column based approach felt like a great framework for complex incidents but rarely for simple ones
For MVP, as we assessed all the feedback, the team aligned on prioritizing single channel view designs over multi channel ones as most of our customer-base was small, medium sized orgs still
As we scale, and sign larger orgs, our designs will be flexible enough to evolve and address multiple channels and highly complex incidents
PRODUCT SUCCESS
After thoroughly testing out with 12 beta orgs, we GA-ed the feature and within a quarter, grew our customer base by ≈20% which we attribute greatly to this feature
Progress Updates and Task Management features continues to show high engagement, task completion rate was up ≈24%
Customer satisfaction (NPS) score was up, with many survey feedback praising our consistent check-ins to address their current and future needs
Leadership wanted us to include more OKRs pertaining to Workstreams and scale
We’re currently designing out and building an onboarding experience to bring more awareness and guidance to set-up and utilize the new feature
POST LAUNCH - AI FEATURE
While monitoring usage of our MVP, we worked on enhancing the UX for our customers to quickly resolve incidents by introducing an AI feature to smartly summarize and create quick, organized postmortems:
WHAT I LEARNED
Having gone through a high-stakes project like this- with constant stakeholder feedback coming from all angles in sporadic times, while required to move fast, I learned that:
Setting more consistent and frequent stakeholder reviews would have helped and sometimes it’s ok to ask design leadership to help round up all the major stakeholders that need to be there
Sitting in even more customer calls with PMs to validate concepts is a good habit to accrue as much strong data points to influence trust from stakeholders
Spending time to persuade engs to work on proof of concept earlier as opposed to waiting on final designs would have been worth it. As Slack <—> WebApp integration was way more complex from a dev perspective than we originally thought
We felt like this feature would be a market differentiator and didn’t invest much time for researching the competitive landscape, but if there are products out there even in a different space that have done something similar and succeeded, an audit would have helped with confirmation of this feature