Validation — Vasl Health

Measured against
the field’s own
instrument.

In 2026 MIT’s McGovern Institute, the Child Mind Institute and Crisis Text Line published the Suicide Risk Lexicon — 50 constructs, 6,910 terms, openly released. We scored VLAP against it on the same corpus and the same split. Every figure on this page carries its run. Nothing here is clinical validation, and nothing VLAP produces is diagnostic: it surfaces language for a clinician’s independent review.

0.8255
Suicide-Attempt Severity — ROC-AUC
Against the Suicide Risk Lexicon’s 0.6489 and a 71-term clinical keyword rule we wrote ourselves at 0.5578. Chunk level, n_pos 307, held-out test split, same treatment for all three.
23 of 47
Signals With No Counterpart
VLAP signals with no match anywhere in the lexicon’s 6,910 terms, across five of our six dimensions — among them the signals for unaliving and catching a fade. Checkable in a minute by anyone who downloads their file.
0.9673
Micro-F1, All 47 Signals
First-person youth check-ins, held-out synthetic text, n = 2,992, run vlap47_20260817T223142Z. This figure stands alone: the clinician-note comparison it used to be paired with was retired on 26 Sept 2026 and no replacement pairing is published.
Synthetic corpus · Non-diagnostic · Not clinical validation
Chapter 01 — What Was Measured

The one decision
that cannot be wrong.

Of the six routing decisions the platform makes, one has no margin: escalate to a clinician now. We validated there first, because it is also the only dimension where an external, clinician-validated, openly published comparator exists.

Against the published lexicon
VLAP severity head — suicide attempt
Chunk level · n_pos 307 · held-out test split
0.8255
Suicide Risk Lexicon — MIT / Child Mind / Crisis Text Line
50 constructs · 6,910 terms · same corpus, same split
0.6489
71-term clinical keyword rule — written by us
Included so the comparison has a floor we control
0.5578
Pre-registered, on real posts
Macro ROC-AUC across four severity classes
14,613 real first-person posts · 1,265 people · three seeds
0.8088
The same keyword rule, same posts
Beaten on all four classes in every seed
0.5074
Attempt, scored per person across all three seeds
Run VLAP-SHA-REAL-001, pre-registered before any model was fit
0.868–0.882
Does it get better with data
PR-AUC gain per 1,000 additional text chunks
95% CI [+0.0083, +0.0234] — excludes zero · run 20260914T002410Z
+0.0159
Training runs behind that slope
25, 50, 75 and 100 percent of available authors
4
Real member language at a quarter of current volume
Already outperforms the entire synthetic-trained model
25%
Chapter 02 — Method and Limits

How to check
all of it.

The comparator is a published artifact with a DOI. The splits are ours, but the instrument is not — which is the point. This is the one claim on the site a reader can falsify without our cooperation.

01
The comparator, named

Suicide Risk Lexicon v1.0 — Low, Ghosh et al., Journal of Psychopathology and Clinical Science, 2026, DOI 10.1037/abn0001152. Released 24 September 2026 by MIT’s McGovern Institute, the Child Mind Institute and Crisis Text Line: 50 constructs, 6,910 terms, clinician-validated, built on roughly 16,000 real crisis conversations. We pulled their released artifact, pinned it by hash, and scored it ourselves rather than reconstructing it.

02
Pre-registered before the fit

VLAP-SHA-REAL-001 was pre-registered before any model was trained — classes, splits, metrics and comparator fixed in advance. Three independent seeds, because one seed is an anecdote. Contamination was checked on author ID and on normalised text hash before scoring, and came back zero on both.

03
What the overlap claim is, and is not

The vocabulary finding is mechanical string overlap and nothing more: 24 of our 47 signals share at least one string with the lexicon, so 23 share none. Whether a given construct and a given signal measure the same underlying thing is a clinical mapping question we have not done and do not assert.

04
The limits, stated

The severity classes are thin at user level, between 9 and 97 positives, and thin classes carry wide intervals. The 0.9673 micro-F1 figure is on synthetic text and is labelled as such everywhere it appears. The real-post corpus is adult writing, not the 15–26 population VLAP serves, so it is supporting evidence rather than the target measurement. None of this is a randomised controlled trial, and none of it is clinical validation.

Non-diagnostic · VLAP never determines, diagnoses, decides or escalates on its own · every output surfaces for a clinician’s independent review
Request Data

See the full
outcomes brief.

The outcomes brief includes detailed methodology, statistical analysis, cohort demographics, disaggregated outcomes by population subgroup, and implementation context. Available to qualified institutional evaluators under NDA.

What's in the brief
Full cohort methodology, statistical analysis, disaggregated outcomes by population subgroup, VLAP signal detection preliminary data, and implementation context for each org type deployed.
IRB study access
The IRB study protocol and preliminary design documentation are available to institutional evaluators under NDA. Contact clinical@vaslhealth.com to request access.
Response time
We respond to outcomes brief requests within one business day. For academic or research institutions, contact research@vaslhealth.com directly.
Related Reading
the research base behind these numbers
The peer-reviewed literature and active IRB study this data builds on.
what these numbers mean for stakeholders
Cost and ROI implications of the retention and symptom data on this page.
the signal layer behind the data
How VLAP's detection informs the clinical escalation figures reported here.
how these figures map to HEDIS measures
PHQ-8 and retention data, reframed against the specific behavioral health measures a payer reports on.