A model breaking in production isn’t just an engineering inconvenience, it’s a live revenue event, and Meta’s ML engineers had no dedicated way to find out why, fast.

Needle in a haystack

How might we help AI/ ML oncalls, model owners and infra maintainers find the “million dollar needle in a haystack.”

I was the solo founding designer for Inference Insights, formerly Hawkeye: a ground-up observability platform for the machine learning systems behind recommendation and ranking products at Meta. I helped take the work from a one-quarter alpha to a broader platform capability over roughly two years, coordinating across on-call engineers, model owners, infrastructure teams, and the Prediction Robustness program.

The opportunity was larger than consolidating dashboards. We needed to turn a fragmented, expert-driven response process into a shared investigation system: fast enough for experienced engineers under incident pressure, legible enough for newer on-calls to use independently, and extensible enough to support needs our small team could not anticipate or build one by one.

In its first two years, Meta reported that Hawkeye facilitated order-of-magnitude improvements in time spent debugging production issues and supported recommendation and ranking models across several products.

Collaboration model

Inference Insights collaboration model
My role

Solo founding product designer across product strategy, research, information architecture, interaction design, prototyping, validation, and design direction.

Team

A core group of roughly three to four UI engineers, backend and infrastructure partners, and a product or technical lead, with periodic UXR consultation and critique from the broader AI Infrastructure design team.

Project scope & phases

Inference Insights project phases
1
Hypothesis validation
  • Identified focused analysis phase to test (Feature upstream traceability)
  • Rolled out alpha to users in under 1 quarter
  • Validated core value
  • Secured leadership backing /Unlocked formal investment
2
Guided investigation
  • Transparent, traceable investigation stages/ visual investigation flow
  • Investigation guidance
  • 1-2 investigators are able to complete investigation in hours instead of days.
3
Automated analyzer integrations
  • Minutes to root-cause for many use-cases
  • Unlocked deeper layers of investigation ie. model parameters/ model internal states (MIS)
  • Complex cases remain reviewable
4
Platform + self-serviceability
  • Formalized into a modular component based platform
  • Self-serviceable/ customizable workflows, user onboarded metrics and analyzers
  • Broader adoption
  • Fewer bespoke customer solutions
Central design challenge

How could we turn 10–15 disconnected debugging tools into one coherent path from production failure to root cause?

Share