PYX LabsBenchmarksPYX-Voice

PYX-Voice

Live

How well frontier models do real employee-listening analysis. Select a model to see how it scored against the expert criteria on a representative task.

Results
Validated results
Leaderboard
Share of tasks passed · 84 tasks
#1muse-spark-1.1
83%
#2gpt-5.6-sol
82%
#3moonshotai/Kimi-K3
82%
#4claude-opus-5
81%
#5gemini-3.8-flash
80%
#6muse-spark-1.3
80%
#7deepseek-ai/DeepSeek-V4-Flash-0731
80%
#8grok-4.6
79%
#9deepseek-ai/DeepSeek-V4-Pro-0813
79%
#10gemini-3.5-flash
76%
#11gpt-6-astra
76%
#12claude-fable-5
76%
#13claude-fable-5-1
75%
#14gpt-5-2025-08-07
71%
#15grok/grok-4
70%
#16claude-sonnet-5
69%
#17gemini-3.1-pro-preview
68%
#18claude-opus-4-8
66%
#19gpt-5-mini-2025-08-07
65%
#20claude-sonnet-4-6
54%
Expert-graded pass rate across 84 tasks.
Rank #1 · In focus
muse-spark-1.1
83%
pass rate
Scored on · employee-experience areas
Performance by employee-experience area
Performance Enablement1.00
Teamwork & Collaboration0.75
Workplace Wellbeing0.89
Recognition & Reward1.00
Growth & Development1.00
Manager Relationship0.67
Change & Innovation0.80
Future / Vision0.64
Engagement0.78
DEIB1.00
Detailed strength/watch-out analysis for this model is pending.

Scores reflect performance on a single representative task, on PYX-Voice's expert-graded 0–100 scale. Model identities are anonymized; see the leaderboard for aggregate results across the benchmark.

A deeper look

How the results break down

The leaderboard above is the summary. These are the same results cut a few different ways — by the type of task, by the underlying capability being tested, and by employee experience topic.

Model performance by task type
Verifiable ResponseSummary
muse-spark-1.1muse-spark-1.1gpt-5.6-solgpt-5.6-solmoonshotai/Kimi-K3moonshotai/Kimi-K3claude-opus-5claude-opus-5gemini-3.8-flashgemini-3.8-flashmuse-spark-1.3muse-spark-1.3deepseek-ai/DeepSeek-V4-Fla…deepseek-ai/DeepSeek-V4-Flash-0731grok-4.6grok-4.6deepseek-ai/DeepSeek-V4-Pro…deepseek-ai/DeepSeek-V4-Pro-0813gemini-3.5-flashgemini-3.5-flashgpt-6-astragpt-6-astraclaude-fable-5claude-fable-5claude-fable-5-1claude-fable-5-1gpt-5-2025-08-07gpt-5-2025-08-07grok/grok-4grok/grok-4claude-sonnet-5claude-sonnet-5gemini-3.1-pro-previewgemini-3.1-pro-previewclaude-opus-4-8claude-opus-4-8gpt-5-mini-2025-08-07gpt-5-mini-2025-08-07claude-sonnet-4-6claude-sonnet-4-60%20%40%60%80%100%
Share of tasks passed, with 95% confidence interval. Verifiable Response (n = 55) tasks have a quantifiable, binary correct/incorrect answer. Summary (n = 29) tasks involve synthesizing open-ended employee feedback into thematic conclusions.
Model by Capability
muse-spark-1.1
0.57
0.79
0.78
0.77
1.00
0.67
0.97
gpt-5.6-sol
0.57
0.86
0.56
0.77
0.80
0.83
0.97
moonshotai/Kimi-K3
0.57
0.85
0.56
0.62
1.00
1.00
0.97
claude-opus-5
0.57
0.79
0.67
0.69
0.80
1.00
0.93
gemini-3.8-flash
0.43
0.71
0.56
0.77
0.80
1.00
0.97
muse-spark-1.3
0.86
0.71
0.44
0.69
0.60
1.00
0.97
deepseek-ai/DeepSeek-V4-Flash-0731
0.57
0.69
0.56
0.69
1.00
1.00
0.93
Synthesis
Summarization
Domain expertise
Multi-step instruction following
Retrieval
Data integrity and anomaly detection
Calculation
1.0
0.8
0.6
0.4
0.2
0.0
Mean score
Proportion of tasks passed per capability, with higher values representing better performance. Each capability is assessed across a minimum of 5 tasks.
Model by Employee Experience Theme
muse-spark-1.1
0.64
0.80
1.00
0.78
0.67
0.75
1.00
0.89
1.00
1.00
gpt-5.6-sol
0.64
0.67
1.00
0.78
0.67
0.50
0.75
0.84
1.00
1.00
moonshotai/Kimi-K3
0.55
0.73
1.00
0.77
0.50
0.75
0.88
0.79
0.83
1.00
claude-opus-5
0.73
0.80
1.00
0.74
0.67
0.75
1.00
0.74
1.00
0.88
gemini-3.8-flash
0.55
0.73
1.00
0.78
0.67
0.75
0.88
0.79
1.00
0.88
muse-spark-1.3
0.73
0.67
1.00
0.74
0.67
0.50
0.88
0.79
0.83
0.88
deepseek-ai/DeepSeek-V4-Flash-0731
0.70
0.67
0.67
0.78
0.80
0.75
0.88
0.74
1.00
1.00
Future / Vision
Change & Innovation
Growth & Development*
Engagement
Manager Relationship
Teamwork & Collaboration*
Recognition & Reward
Workplace Wellbeing
Performance Enablement
DEIB
1.0
0.8
0.6
0.4
0.2
0.0
Mean score
Performance by employee experience theme (People Insights Model). Higher is stronger. *Indicates factors with lower coverage; results should be interpreted directionally.