The problem


The Farmer's Dog is a subscription pet food company with $800M+ in annual revenue and a 400+ person Customer Care team handling over 50,000 customer contacts a week. Quality Assurance (how TFD evaluates those conversations and coaches advisors to improve) is what makes "Iconic Care" consistent, not aspirational.

But less than 2% of conversations were being reviewed. The existing tool, EchoAI, cost ~$500K/year and was rated 3.48/5, the lowest of any internal platform. It lacked TFD's domain knowledge, coaching was reactive, and performance gaps went undetected. The company had no real control over how quality was being measured.

The goal: scale to 75%+ conversation coverage via AI, replace EchoAI, save ~$200K/year, and build direct control over quality measurement tailored to TFD.



My role


Lead Product Designer. I owned the product experience end-to-end — from 0 to shipped, in a space with more questions than answers. I worked directly with Engineering, Legal, Customer Care Ops, and Data throughout.












Challenges

Several constraints hit at once




I joined week two with no PM, an unfamiliar domain, and a newly formed team. I combined onboarding with early information-gathering, getting to know the people while learning how QA actually worked in practice.

Four pressures were already in play:

  1. Legal. AI evaluation of employee conversations is genuinely ambiguous territory, especially in California. The distinction between "coaching" and "surveillance" isn't semantic. It's legally significant.
  2. Technical. LLM quality, consistency, and cost were all unsettled. Real-time vs. batch processing had tradeoffs across accuracy, user experience, and budget.
  3. Users. AI sentiment ranged from skeptical to anxious. Trust in EchoAI was already broken. We weren't starting from neutral.
  4. Business. EchoAI's contract was ending September 2025. The clock was running












May–Oct 2025

Phase 1: Building the foundation





The core idea


Phase 1 launched with manual grading only. No AI. The logic was straightforward: get people onboarded first, build trust with the tool itself, understand how users actually work, then bring AI in once the foundation is solid. Shipping AI on day one would have meant building on unverified assumptions, with users who had already lost confidence in the last platform.









Design moved fast


Low-fi wires and user flows, validated with early research to define core workflows and feature priorities. Then mid-fi prototypes, usability tested on the core grading flow. Then hi-fi mocks, full eng handoff, and demo prototypes for CC training before launch.

Usability testing (n=5) revealed things we wouldn't have caught otherwise. Decimal scores (like 16.67) were confusing and felt demotivating. Editing one score wiped all previously entered scores and notes, the single biggest friction point. The scoring dropdown hid options behind too many clicks.

I iterated on all of it. Decimal scores were removed in favor of plain-language ratings with color coding. One-click scoring with exposed options replaced the dropdown. These felt like small fixes, but they're the kind of details that told users we'd actually thought about their workflow. Ease of use landed at 4.8/5, confidence at 4.6/5.





Strategic decisions I made under constraint


I chose Material UI over TFD's internal design system. This was a deliberate trade-off, aligned with the eng lead. Advisors with color blindness had flagged contrast issues on our existing system. MUI is robust, well-documented, and AI tools generate its code more accurately, which mattered with a backend-focused eng team, 0.5 of a designer, and a tight timeline. Another team was already using it, so it reduced ramp-up. Not the obvious call, but the right one.

I kept the UI intentionally scaled back. Clean, functional layouts focused on information hierarchy and usability. High-volume QA work needs clarity and speed, not visual complexity. This also meant eng could ship features without deep designer involvement, so we moved faster with limited resources.

When Figma MCP launched in June 2025, I pushed the team to experiment with it immediately. AI-human optimized annotations doubled our design handoff velocity while maintaining quality. We ran an internal Lunch & Learn afterward, and several other teams at TFD adopted it into their workflows.









Phase 1 / Key Flow 1 / Select Agent to Review












Phase 1 / Key Flow 2 / Grade Manually










Phase 1 / Key Flow 3 / Edit Scores








Phase 1 landed well


Satisfaction jumped from 3.48 to 4.05/5. Graded conversations scaled from a trickle to hundreds per week. Users trusted the platform. Fast development, high implementation fidelity, flawless execution. We had a strong foundation.















Phase 1 → Phase 2 Transition

53% preferred little or no AI






The research bridge


Before designing Phase 2, I recommended running a survey to understand what we were actually walking into. 300+ respondents. The questions I really wanted answered: how do people feel about the platform now, what's their honest AI sentiment, and what do they want automated vs. kept human.

The platform sentiment was good. But the AI numbers stopped me.




The survey came back


~53% preferred little or no AI in QA.

95% cited accuracy and contextual judgment as their top concerns. Graders told us finding gradable conversations was their biggest time sink. Filters reset every session, no way to preview before opening. Advisors told us feedback arrived weeks after the conversation, too late to act on. And the QA team wanted cumulative trends, not one-off evaluations.

People were open to AI for faster coverage and trend-spotting. But not for replacing human grading. Not for making decisions about their performance. Not yet.

This was the real design problem. We were building an AI-powered evaluation system for people who mostly didn't want AI evaluating them. The trust we'd built in Phase 1 was real, but the next phase could break it if we got the approach wrong.














Oct 2025–Apr 2026

Phase 2: Integrating AI






Challenges


I started exploring custom components and establishing AI interaction patterns and on-brand design language. The core design challenge wasn't how to display AI scores. It was how to introduce AI in a way that felt genuinely on the user's side.

Three tensions shaped every decision:

  1. AI automation needs to run at scale to deliver value, but employees need agency and transparency, not surveillance. Human review still matters. AI can't replace context.
  2. California has the strictest employee monitoring laws in the US. "Coaching" vs. "surveillance" is a legally significant distinction. Results need to be unbiased.
  3. LLM quality and consistency are real concerns. Real-time vs. delayed vs. batch processing each have tradeoffs. And you have to design for the cases where AI output is low-quality or off-base.




Principles I designed around


  1. Transparency, not automation. Auto review and manual review coexist side by side. Users see both. Human judgment is always the final authority. I reframed the language: "AI grade" became "Auto review," "Human grade" became "Manual review." Small shift, significant difference in how it feels.
  2. Minimal intervention. Users can toggle AI review off entirely for manual-only sessions or in-person coaching. "How auto-review works" is always one click away. The feedback mechanism is framed as "help us improve coaching," not "report a problem."
  3. Legal compliance as design input. I worked directly with Legal throughout. AI scores are explicitly not used for performance grading or disciplinary decisions, only for coaching. That constraint is built into the UI itself: clear legal disclosure, persistent documentation, usage guardrails. Not buried in a terms page.








Phase 2 / Key Flow 1 / View & Filter Conversations










Phase 2 / Key Flow 2 / View Graded Rubrics












Phase 2 / Key Flow 1 / Live Grade Manually









What we decided not to build


Some of the most important product decisions were cuts.

  1. I didn't show AI scores to Advisors directly. We're still training the system. Showing someone an AI evaluation of their own performance before we're confident in its accuracy isn't transparency. It's noise that erodes trust. We'll get there, but not before the system earns it.
  2. I didn't build AI flagging. We considered it. AI would surface conversations that need attention. But the QA team itself wasn't sure what should be flagged. Building a smart feature on top of uncertain behavior would have created more confusion than value. We need more data first.
  3. I didn't build an editable draft UI for AI-generated notes. Users asked for it. But the scope was enormous: editing, saving, version control. And the real insight was simpler: only the QA head controls the rubric versions. One person on the backend. It's not a UI problem.

Each cut was the same judgment in a different form: we knew what we didn't know, and designed accordingly.















Outcomes

  • Cost: Replaced EchoAI, eliminating ~$200K/year in licensing costs.
  • Coverage: From ~3% to 75%+ conversations reviewable with AI.
  • Satisfaction: 4.05/5 vs. EchoAI's 3.48/5 (+0.57 improvement), 337 respondents.
  • Feedback speed: From 5 days average to target <1 day with Phase 2.
  • Adoption: 87% of survey respondents described EvalPal as a "coaching reinforcement" tool, matching the design intent exactly.
  • Ripple effect: Figma MCP workflow adopted by multiple teams after internal Lunch & Learn.




The unexpected proof point

Legal bias audit passed. No significant scoring differences across race, ethnicity, gender, or age. EvalPal actually outperformed human grading in minimizing demographic score differences. The constraint that felt like the biggest risk turned into the strongest proof point.










Reflection


Designing AI tools for people who didn't ask for AI is a different problem than building for users who opted in.

The people using EvalPal don't get to choose whether their conversations are evaluated. That changes what good design means. It's not enough to be usable. The product has to be genuinely on their side. Every transparency decision, every language choice, every feature we cut was really an answer to the same question: does this serve the person being evaluated, or just the system doing the evaluating?

The constraints helped more than I expected. Legal ambiguity forced us to build transparency into the product itself, not as an afterthought. Limited design resources forced a cleaner, more functional UI. A backend-focused eng team made MUI the right call over custom components. The trust-first phasing meant Phase 2 didn't have to fight for adoption. It was invited in.

What we didn't build mattered as much as what we did. Not showing AI scores to Advisors. Not building flagging without behavioral data. Not overengineering a draft editor that one person could handle in the backend. Each cut protected the thing that made the product work: the belief that the system was being honest with you.

I keep coming back to this: the hardest part of designing an AI product isn't the AI. It's the relationship between the AI and the person it's pointed at. Get that wrong, and nothing else matters.


@Xiaoyu Liu