Signal · AI Labs

Measure human and AI performance — apart, and together.

Models are ranked against other models. Almost nothing measures whether a person actually performed better with one. Signal runs the comparison inside real gameplay: people unaided, AI alone, and people amplified by AI — and reports the lift.

Research, white papers and the live benchmark live at skillprint.co/benchmark.
Human+AI benchmarks

Measure the lift, not just the model.

A leaderboard position says a model is good at the benchmark. It says nothing about whether a person got further with it in their hands. Signal measures the difference, and the difference is the thing worth optimising.

Three arms, one comparison

Human+AI against AI alone, and against people working unaided, run blind inside real games rather than on a static test set.

Human+AIAI aloneUnaided

11 frontier models ranked

Gemini, Claude, GPT and open weight models, scored on goal attainment, cost and efficacy by one neutral scorer.

Goal attainmentCostEfficacy

Statistics that hold up

Bayesian averaged ratings, 95% confidence intervals and Welch's t test. A live benchmark keeps us precise.

Bayesian95% CIWelch's t test
What it answers

Where do humans and AI perform best together?

The benchmark exists to answer questions a model only leaderboard structurally cannot.

01
Does this model actually improve the person using it?Measured against the same person's unaided baseline, in the same task, on the same day.
02
Which model amplifies which kind of person?Lift is not uniform. A model that helps a cautious planner may do nothing for a fast improviser.
03
Where is the human still better alone?The control arm runs with no AI in the loop, so the cases where assistance hurts are visible rather than assumed away.
04
What does the lift cost?Goal attainment scored beside token cost and latency, so efficacy is judged against what it took to get there.
The leaderboard

See how models score against human play.

Skillprint Labs ranks frontier models from real gameplay sessions. Mood alignment is one of the boards. The benchmark is open, and it updates as sessions land.

The Skillprint Labs AI benchmark: eleven evaluated models, 2,846 benchmark play sessions, and a bar chart ranking models by average mood alignment rating
The dataset

Every session is a labelled reasoning example.

Training data is overwhelmingly what people produced once they had finished thinking. Gameplay captures the process instead: the decisions, the recoveries and the adaptations that produced the output.

Talk to research
01

Real world relevance

Gameplay maps to how people actually reason and decide, not to how they describe it afterwards.

relevance
02

Structured labels at scale

Every session returns schema constrained output against one ontology, so it is comparable across games.

labels
03

Measurable human+AI lift

Benchmarks show which systems genuinely improve people, and by how much.

lift
Work with us

Open Skillprint Labs.

Research access, dataset scope, benchmark participation and governance. Tell us what you want to measure.

A SkillprintTwelve of the cognitive skills and moods a Skillprint holds, from problem solving and memory through to focus and collaboration. Each one runs as a strand into the Skillprint mark at the centre, where they combine into a single profile.Problem SolvingMemorySpeedAccuracyPattern RecognitionSpatial AwarenessLogicCreativityInnovateRelaxFocusCollaborate