In 2026 MIT’s McGovern Institute, the Child Mind Institute and Crisis Text Line published the Suicide Risk Lexicon — 50 constructs, 6,910 terms, openly released. We scored VLAP against it on the same corpus and the same split. Every figure on this page carries its run. Nothing here is clinical validation, and nothing VLAP produces is diagnostic: it surfaces language for a clinician’s independent review.
vlap47_20260817T223142Z. This figure stands alone: the clinician-note comparison it used to be paired with was retired on 26 Sept 2026 and no replacement pairing is published.Of the six routing decisions the platform makes, one has no margin: escalate to a clinician now. We validated there first, because it is also the only dimension where an external, clinician-validated, openly published comparator exists.
The comparator is a published artifact with a DOI. The splits are ours, but the instrument is not — which is the point. This is the one claim on the site a reader can falsify without our cooperation.
Suicide Risk Lexicon v1.0 — Low, Ghosh et al., Journal of Psychopathology and Clinical Science, 2026, DOI 10.1037/abn0001152. Released 24 September 2026 by MIT’s McGovern Institute, the Child Mind Institute and Crisis Text Line: 50 constructs, 6,910 terms, clinician-validated, built on roughly 16,000 real crisis conversations. We pulled their released artifact, pinned it by hash, and scored it ourselves rather than reconstructing it.
VLAP-SHA-REAL-001 was pre-registered before any model was trained — classes, splits, metrics and comparator fixed in advance. Three independent seeds, because one seed is an anecdote. Contamination was checked on author ID and on normalised text hash before scoring, and came back zero on both.
The vocabulary finding is mechanical string overlap and nothing more: 24 of our 47 signals share at least one string with the lexicon, so 23 share none. Whether a given construct and a given signal measure the same underlying thing is a clinical mapping question we have not done and do not assert.
The severity classes are thin at user level, between 9 and 97 positives, and thin classes carry wide intervals. The 0.9673 micro-F1 figure is on synthetic text and is labelled as such everywhere it appears. The real-post corpus is adult writing, not the 15–26 population VLAP serves, so it is supporting evidence rather than the target measurement. None of this is a randomised controlled trial, and none of it is clinical validation.
The outcomes brief includes detailed methodology, statistical analysis, cohort demographics, disaggregated outcomes by population subgroup, and implementation context. Available to qualified institutional evaluators under NDA.