Vasl Health › Insights › Research
Research

The Data That
Never Existed

A corpus pairing African American Language with behavioral-health distress labels did not exist in any public repository. We checked. Then we built it.

R
Rodney Bell
Founder & CEO, Vasl Health
September 2026
7 min read
Research

Black youth suicide rose 29.4% between 2018 and 2023. White youth suicide fell 14.8% over the same period.1 Those two numbers sat next to each other in a CDC report, and the detection systems deployed in schools and clinics during those five years were trained and validated on everyone except the population where the rate was climbing fastest.

The instinct is to call this an awareness problem — more training, better outreach, a poster in the counselor’s office. It is not an awareness problem. Awareness does not change what a language model was trained on. The gap is data. Specifically: a corpus pairing African American Language (AAL) with behavioral-health distress labels, annotated by people who speak the language and trained to recognize clinical risk inside it. We looked for that corpus before we built one. It does not exist in any public repository. We checked systematically, not casually — searched the standard NLP dataset registries, the academic corpora indexed for clinical and social-science research, and the commercial vendors who license behavioral-health training data. Nothing pairs the two.

Why the Gap Is Structural

Standard sentiment and risk-detection models are trained on internet text that skews heavily toward Standard American English. A young person who codes-witches into AAL under stress, or who expresses distress through register shift rather than explicit language, is invisible to a model that has never seen that pattern labeled as anything at all — let alone labeled as risk. The model does not fail loudly. It fails silently, by never flagging the signal in the first place.

This is not a bias that gets fixed by adding more examples to an existing dataset. There was no existing dataset with the right label structure to add examples to. Someone had to define the taxonomy first — VLAP’s 47 signals across 6 categories — and then annotate real, community-sourced language against it, with clinicians adjudicating what counts as risk and what counts as culturally normal register. That is a multi-year construction project, not a data-cleaning task.

“The data to fix that cannot be bought, licensed, or downloaded. It has to be built with the community, from inside the language.”
Rodney Bell — Founder & CEO, Vasl Health

What “Community-Annotated” Actually Means

It means the people labeling the data speak the language natively and were trained on the clinical taxonomy, not the reverse. A generic crowdworker platform cannot annotate AAL distress markers accurately — the same phrase can be an idiom, a joke between friends, or a genuine flag, and reading that correctly requires cultural fluency a general-purpose labeling pipeline does not have. Vasl built its annotation process around that requirement from the start, which is a large part of why the resulting library runs to more than 2,000 tokens and phrases mapped against clinical categories rather than a generic sentiment scale.

The output is the first clinical-language platform built on a corpus like this. Not the first to claim cultural sensitivity as a feature — the first to have the annotated data underneath the claim.

One Layer, Many Modules

The architecture is intentionally not “an AAL model.” It is one intelligence layer with a language module for every community it serves, built with that community. African American Language is the first module because the CDC data made the case first and most urgently. Rural and Appalachian youth are next — the same community-annotated method, the same clinician-in-the-loop process, applied to counties where the behavioral-health workforce gap is not about cultural mismatch but about there being no workforce at all.

Every module that gets added follows the same rule: VLAP surfaces distress signals for a clinician’s independent review. It does not monitor continuously in the background, it does not escalate on its own, and it does not diagnose. A clinician reads what the model surfaces and decides what happens next. That constraint is not a limitation we are working around — it is the reason the taxonomy had to be built by clinicians in the first place, and the reason the corpus took as long as it did.

1 — CDC, MMWR 74(35);550–553, “Differences in Suicide Rates by Race, Ethnicity and Age Group, 2018–2023” (Sept. 2025)

See Vasl Health
in action.

A 45-minute call with an implementation lead — not a sales pitch. We come prepared with context specific to your organization and population.

Request a Demo →
Related Reading
the missing infrastructure
Why the community-annotated corpus behind VLAP had to be built, not licensed — and what comes after African American Language.
the model built for this population
What training a language model directly on underserved communities' speech patterns actually requires.