PYX LabsBlogWhat Frontier AI Gets Wrong About Employee Feedback: Inside the First PYX Labs Webcast

What Frontier AI Gets Wrong About Employee Feedback: Inside the First PYX Labs Webcast

What Frontier AI Gets Wrong About Employee Feedback: Inside the First PYX Labs Webcast

Today we hosted the first PYX Labs webcast, "Sounds Right, Isn't: Where Frontier AI Models Fall Short on Employee Feedback." Joe Freed (Founder, PYX Labs) and Sarah Foster (Director, PYX Labs) were joined by Melissa Valentine, PhD — a Stanford professor, Senior Fellow at the Stanford Institute for Human-Centered AI (HAI), and founder of Stanford's new AI and Organizations Lab — to walk through what our first benchmark, PYX-Voice, revealed about how well today's frontier models actually understand employee feedback.

If you missed it, here's the short version, plus the full recording.

The question at the center

Companies are already leaning on AI to interpret how their people feel at work — pasting survey results into ChatGPT, drafting action plans off a listening pulse, asking a model to summarize what employees really mean. The question no one had answered: how good is AI at that, actually?

"We wanted to understand one thing clearly," Joe said. "How well do AI models actually understand employee feedback?" Evaluating AI on objective, pass/fail tasks is relatively straightforward, and it's understandably where earlier benchmarks focused. The harder, more subjective work of the workplace is different: reading nuanced employee feedback and turning it into something a leader can reliably act on. That's where our work is focused now.

Why the evaluation criteria are the real frontier

Melissa framed the bigger picture. Her lab studies how AI is changing the workplace, and a recurring theme in AI evaluation is benchmarks built on real-world tasks — Mercor and Upwork both released sets showing frontier models struggle with actual professional work, such as investment banking, management consulting, and corporate law. Digging into why, she landed on something foundational: the rubrics. Defining what "good" looks like, and having human experts grade against it, is the hard, decisive layer of any benchmark.

"Subjectivity is a neutral word — it just means a human is making a judgment," she said. "And that's an opportunity for a lot of values, a lot of expertise, a lot of leadership, and a lot of strategy." How a company defines and evaluates AI performance, she argued, is where its values and strategy actually live. That's what made PYX-Voice distinctive to her: the time spent formalizing what a good answer looks like in employee experience.

How we built PYX-Voice

Sarah, an I-O psychologist who leads the lab, walked through the build:

  • Started with a job analysis of Perceptyx's employee-listening consultants, then narrowed the benchmark to two parts of that job: analysis and interpretation of employee-listening data.
  • Used real data. 17,500 real employees' responses, de-identified and scrubbed, organized into three hypothetical organizations.
  • Wrote the tasks and rubrics. Each task paired with a gold-standard expert response and criteria defining what "good" looked like, across three task types on a spectrum of subjectivity: verifiable response (objective), summary (moderately subjective), and executive narrative (highly subjective, action-oriented).
  • Graded the graders. Two rounds of human expert grading were used to train an LLM judge — and to confirm humans agreed before trusting a model to score. The most subjective "executive narrative" tasks produced inter-rater reliability problems, so 21 were set aside for now.
  • Result: 84 final tasks and 200+ discrete criteria, used to rank seven frontier models from OpenAI, Google, Anthropic, and xAI.

As Sarah noted, quoting fellow advisor Ethan Burris of UT Austin: "If you put the same employee-listening task in front of 20 different I-O psychologists, they'd have completely different output — and a completely different definition of what good looks like." Aligning experts on that definition is the core challenge the lab is built to solve.

What we found

  • Reasoning power didn't predict performance. Gemini-3.5-flash — fast and cost-efficient, running at low reasoning effort — was the top model. Models positioned for heavier reasoning, like Claude, scored lower. The top model led around 76%; the field ranged down to about 54%.
  • Stronger on the objective, weaker on the subjective. Models did better on verifiable-response tasks than on summary tasks — and, to our surprise, handled quantitative and statistical work reasonably well, despite the "AI is bad at math" assumption.
  • Critical errors were rare but real. Hallucinations and overclaims showed up occasionally. In a domain with this low a tolerance for error, even rare failures give pause.
  • Synthesis was the weakest capability. Pulling disparate data together into a coherent interpretation — connecting the dots — is exactly the high-value work practitioners want help with, and it's where models struggled most.
  • The more human the topic, the harder it got. Using the Perceptyx People Insights Model, models did best on clear, consistently worded topics like Performance Enablement, and worst on nuanced, context-specific ones like Change & Innovation and Future/Vision. Every model struggled the same way on the most human themes.

The takeaway: not good or bad, but where

"The main takeaway isn't that models are good or bad," Joe said. "It's that they're weaker or stronger in some places." That makes PYX-Voice a map for where a human needs to stay in the loop, and where prompting, frameworks, and expertise add the most value. Crucially, these are the base frontier models — the gap between a raw model and a purpose-built, expert-grounded system is exactly the value the field can bring.

The failure mode that stuck with everyone: on the most valuable, action-oriented work, models were confident but unfounded. "Pie in the sky, but very confident," as Joe put it — recommendations that sounded authoritative but weren't grounded in the data or the science.

What's next

The next iteration focuses on the hardest, highest-value part: actionability — the "so what, now what" of employee listening. That means breaking subjective tasks into smaller, more discrete components, sharpening the criteria, and adding behavioral-science "sturdiness" checks so recommendations are grounded in evidence rather than management fads. Melissa called it "building out loud" — putting a point of view on record and letting people poke at it. In a field with little public discussion of how AI actually performs on this work, that's the point.

[▶ Watch the full webcast above] for the complete conversation, including the live discussion of criteria, values, and where AI evaluation goes from here.


Frequently asked questions

What is PYX-Voice?
PYX-Voice is the first benchmark from PYX Labs, measuring how well frontier AI models perform real employee-listening work — analyzing survey data and interpreting open-ended feedback — graded against criteria written by I-O psychologists. It evaluated seven frontier models across 84 tasks and 200+ criteria.

Which AI model was best at understanding employee feedback?
Gemini-3.5-flash ranked highest overall (around 76%), and it did so even at low reasoning effort; models built for heavier reasoning, like Claude, didn't automatically do better. But as Joe Freed pointed out during the webcast, the ranking is almost beside the point. The real finding isn't that any one model is good or bad; it's that every model was strong on some topics and weak on others, and all of them struggled with the same nuanced, uniquely human themes. So the useful question isn't which model wins, it's knowing where a given model can be trusted and where it can't.

Is AI good at analyzing employee survey data?
It depends on the task. Models handled objective, quantitative work reasonably well, but were less reliable on subjective interpretation — reading nuanced, emotional feedback and synthesizing it into coherent themes. That interpretive work is where they fell short and where human expertise still matters most.

Should companies use AI for employee feedback?
Yes, with care. The webcast's message wasn't "don't use AI" — it was know where it's reliable and where it needs a human in the loop, better prompting, or expert-built guardrails before it informs a real decision about people.

Who ran the study and who advised it?
PYX Labs, a research initiative sponsored by Perceptyx, built PYX-Voice. It's advised by Melissa Valentine, PhD (Stanford / Stanford HAI) and Ethan Burris, PhD (UT Austin).

Where can I read the full benchmark report?
The full benchmark report, along with methodology and per-model scores, can be found at https://www.pyxlabs.ai/research.

← Back to all posts