Research notes from the lab.
Method notes, findings, and field observations.

What Frontier AI Gets Wrong About Employee Feedback: Inside the First PYX Labs Webcast
In our first webcast, Joe Freed and Sarah Foster sat down with Stanford's Melissa Valentine to unpack what PYX-Voice found when we tested seven frontier AI models on real employee-listening work. The short version: the models are more fluent than they are reliable, reasoning power didn't predict performance, and the harder and more human the feedback, the more they struggled. Here's the recap, plus the full recording.
Read →
AI Can Summarize Employee Feedback. A New Benchmark Shows It Doesn't Always Understand It.
PYX-Voice, the industry's first benchmark for how AI understands employee feedback, finds it often misses the meaning behind complex workplace experiences.
Read →
Introducing PYX Labs: How Good Is AI, Really, at Understanding People at Work?
AI already shapes how organizations read employee feedback. PYX-Voice, our first benchmark, found frontier models far stronger on numbers than interpretation.
Read →