Where the number comes from
+0.8348
Agreement between the human rater and GPT-4, scoring the same 271 occupations
Both columns are per-occupation exposure ratings of the same 271 jobs — one by a human rater, one by GPT-4. They correlate +0.8348, 95% CI +0.7947 to +0.8676, with a mean absolute gap of 0.077 and a maximum of 0.643.
GPT-4 rates these far more exposed than the human did
| occupation | human | gpt-4 | gap |
|---|---|---|---|
| Medical transcriptionists | 0.232 | 0.875 | +0.643 |
| Bookkeeping, accounting, and auditing clerks | 0.314 | 0.802 | +0.488 |
| Court reporters and simultaneous captioners | 0.521 | 0.958 | +0.437 |
| Music directors and composers | 0.281 | 0.711 | +0.430 |
| Computer hardware engineers | 0.318 | 0.727 | +0.409 |
| Musicians and singers | 0.133 | 0.408 | +0.275 |
and these far less
| occupation | human | gpt-4 | gap |
|---|---|---|---|
| Concierges | 0.700 | 0.467 | −0.233 |
| Fitness trainers and instructors | 0.235 | 0.015 | −0.220 |
| Survey researchers | 0.844 | 0.625 | −0.219 |
| Agricultural and food scientists | 0.658 | 0.456 | −0.202 |
| Childcare workers | 0.340 | 0.141 | −0.199 |
| Public relations specialists | 0.788 | 0.591 | −0.197 |
Read the two lists as pairs. Transcription, bookkeeping, court reporting, composition — work that arrives as text or symbols. Childcare, fitness instruction, concierge work — work that requires being in the room. Writers and authors sits at 0.774 human against 0.877 GPT-4, a gap of +0.103, ranked 47 of 271: GPT-4 thinks writing is even more exposed than the human rater does.
not independent These raters are not independent of each other in any strong sense — both are scoring the same occupation descriptions, and the human rater may have had model output available. Treat the agreement as a consistency check, not as replication.
the pattern is the point A correlation of +0.8348 between two raters means the disagreements are a small minority of cases. What makes them worth showing is that they are not scattered — they sort cleanly by medium of output, the same split the ability layer produces.
direction unknown Nothing here says which rater is right. It is entirely possible that GPT-4 is correct about transcription and the human rater is correct about concierges. The claim is only that the two disagree along one axis, and that the axis is not cognitive difficulty.
