Where the number comes from
14
The line the source file draws across the questionnaire — and the check that it really is that line
The questionnaire total runs from 0 to 51 with no gap anywhere along it. There is no place in the distribution where “well” becomes “unwell”; the line is drawn, not found.
| the line, and what it flags | |
|---|---|
| cutoff used by the file | 14 |
| teenagers at or above it | 787 of 4,810 |
| share flagged | 16.4% |
| questionnaire range | 0 – 51 |
| median total | 5 |
| mean total | 7.27 |
the cutoff was verified, not assumed
The source ships a binary depressed column without stating what produced it. So the pipeline tries every cutoff in range and reports how well each reproduces that column across all 4,810 rows. Exactly one is perfect, and the build asserts it — if a future version of the dataset moved the line, the pipeline would fail rather than quietly publish a wrong claim.
| cutoff | agreement with the file's own column |
|---|---|
| bdi_total ≥ 11 | 93.16% |
| bdi_total ≥ 12 | 95.72% |
| bdi_total ≥ 13 | 97.98% |
| bdi_total ≥ 14 | 100.00% |
| bdi_total ≥ 15 | 97.78% |
| bdi_total ≥ 16 | 96.49% |
| bdi_total ≥ 17 | 94.80% |
| bdi_total ≥ 18 | 93.45% |
| bdi_total ≥ 19 | 92.77% |
| bdi_total ≥ 20 | 91.83% |
| bdi_total ≥ 21 | 90.52% |
bdi_total ≥ 14 reproduces the column exactly. The next best, 13, gets 97.98% — close, and wrong.
