Chapter 01 — The Gap

Standard AI was
trained on the
wrong language.
We fixed that.

Every major NLP model in clinical use was trained predominantly on majority-White internet text. That corpus doesn't contain AAVE. It doesn't contain queer vernacular. It doesn't contain the coded language underserved youth use to describe pain. The gap between what the model was trained on and who it's supposed to serve is not a fine-tuning problem. It's an architecture problem. VLAP — our clinical language model, a fine-tuned BERT-family transformer — is the answer.

This page covers the model at a glance. For the full architecture, training methodology, and security documentation, see the Technology deep-dive.

Underlying Model Architecture
VLAP
A fine-tuned BERT-family transformer model trained on culturally-specific mental health language corpora drawn from Black, Latino, LGBTQ+, and youth in urban and rural communities. The foundational layer of the Vasl Language Analysis Platform.
90%
Sensitivity on high-distress
signal detection
47
Clinical signals
across 6 categories
2,400+
AAVE & youth vernacular
tokens in vocabulary
Chapter 02 — The Model

Not fine-tuned
on culture.
Trained from
inside it.

VLAP begins with a fine-tuned BERT-family architecture base and extends it with training samples drawn from culturally-specific mental health language — AAVE, queer vernacular, code-switching, trauma-adjacent idiom, and youth-specific coded language. This is not a demographic filter applied to a standard model. The cultural specificity is in the weights.

"A model trained on the wrong language will always be measuring the wrong thing. No fine-tuning fixes a training set."
Training Corpus — VLAP V1
Base architecture
BERT-family transformer — fine-tuned on domain-specific corpora with culturally extended vocabulary layer
Training samples
Utterances drawn from Black, Latino, LGBTQ+, and youth in urban and rural communities mental health language corpora — curated, annotated, and reviewed by community-embedded clinical staff
Vocabulary extension
2,400+ tokens added to base vocabulary: AAVE constructions, queer vernacular, youth coded language (including euphemisms for self-harm ideation), minimization patterns, and cultural idiom
Annotation method
Signal labels assigned by licensed clinicians with cultural competency training — not crowdsourced, not automated. Inter-annotator agreement validated above clinical threshold before inclusion
Bias monitoring
False positive and false negative rates tracked by demographic subgroup in every inference batch. Systematic disparities trigger mandatory model review — not a reporting metric, an operational gate
Validation status
Active IRB study in partnership with a university research partner validating clinical signal accuracy in production deployment. Study ongoing — results to be published upon completion

What Standard NLP Misses

Pattern 01 — AAVE Hopelessness Constructions
"idk why i even try anymore tbh, can't keep doing this no more fr."
Classified as informal or low-confidence text. No distress signal fired. Member not flagged.
VLAP reads: AAVE futility construction (HOP-03) + AAVE minimization construction (CCM-02). Signal confidence high. Surfaced to clinician with interpretive context.
Pattern 02 — Coded Self-Harm Language
"been thinking about unaliving lately ngl."
Token "unaliving" absent from vocabulary. Statement classified as neutral or ambiguous. No signal generated.
VLAP reads: unaliving and derivatives (SHA-02). High-risk signal — surfaced immediately to clinical supervisor for human review per escalation rules.
Pattern 03 — Pre-Disclosure Minimization
"it's not that deep but lowkey been struggling since school started."
Minimization phrase ("not that deep") reduces overall signal weight. Statement classified as moderate-concern, often deprioritized.
VLAP reads: pre-disclosure minimization (CCM-04) increases contextual weight on what follows. The minimization is the signal — not evidence against one.
Pattern 04 — Performative Positivity
"everything is fine now actually, I'm so over it lol."
Positive sentiment classification. Prior distress signal weight reduced. Member status improved in automated view.
VLAP reads: sudden positive shift following distress as high-risk farewell pattern (HOP-07). Escalation to clinical review — not a resolution signal.
Chapter 03 — Signal Taxonomy

47 signals.
6 categories.
Zero diagnostic
outputs.

The VLAP Distress Signal Taxonomy (v1.0) defines every signal VLAP is trained to detect. Each signal has a clinical definition, a community-language annotation, and a risk-level weighting. None produce a diagnosis. All surface as interpretive context for a licensed clinician — who decides what to do with them. Full clinical notes, escalation rules, and validation status for every signal are documented in the Technology deep-dive.

HOP-01 – HOP-09
9
Hopelessness & Futility

The Beck Hopelessness Scale and Columbia Suicide Severity Rating Scale both identify hopelessness as the strongest independent predictor of suicidal ideation. Generic NLP models perform adequately on Standard American English expressions of hopelessness (HOP-01). VLAP’s differentiation is in detecting culturally coded and vernacular expressions of futility (HOP-02 through HOP-09) that generic models systematically misclassify as frustration or hyperbole.

HOP-01HighExplicit future negation (“there's no point anymore”)
HOP-02ModerateAmbiguous future self-erasure (“I won't be here when that happens”)
HOP-03HighAAVE futility construction (“can't keep doing this no more fr”)
HOP-04VariableCoded exhaustion (“tired of waking up and doing this again”)
HOP-05LowTemporal foreclosure (“stopped making plans”)
HOP-06ModerateThird-person self-distancing (“that's not me anymore”)
HOP-07HighBurden narrative (“everyone would honestly be better off”)
HOP-08HighQueer futility narrative (“it doesn't get better for people like me”)
HOP-09ModerateImmigrant / first-gen foreclosure (“all that sacrifice and for what”)
Signal Example
"I just need to get through this week, fr."
HOP-04 — Coded exhaustion. “Fr” (for real) is an authenticity marker, not casual speech.
"It doesn't get better for people like me."
HOP-08 — Queer futility narrative. Carries the weight of observed community trauma, not just personal pessimism.
SHA-01 – SHA-08
8
Coded Suicidal Ideation

The SHA category captures both explicit and coded expressions of suicidal ideation. Generic models perform well on explicit ideation (SHA-01) and poorly on all other signals. SHA-02 through SHA-08 represent the core of VLAP’s clinical differentiation — these are the expressions young people actually use, developed specifically to evade platform content moderation and the perceived stigma of explicit disclosure.

SHA-01HighExplicit suicidal ideation (“I'm thinking about suicide”)
SHA-02HighUnaliving and derivatives (“been thinking about unaliving”)
SHA-03ModerateSleep-death conflation (“I just want to go to sleep and not wake up”)
SHA-04ModerateFantasy of absence (“what would happen if I disappeared”)
SHA-05HighMethod inquiry with minimization wrapper (“just curious how many pills”)
SHA-06HighGoodbye behavior references (“made sure everything was in order”)
SHA-07HighDate markers with inverted framing (“I just need to get through [date]”)
SHA-08HighCommunity-specific death euphemisms (“catching a fade fr”)
Signal Example
"been thinking about unaliving lately ngl."
SHA-02 — Coded ideation. “Ngl” is a sincerity marker, not minimization.
"not that I'd do it but what does it feel like when..."
SHA-05 — Method inquiry with minimization wrapper. Always triggers human review regardless of stated intent.
CCM-01 – CCM-08
8
Minimization & Disclosure Suppression

CCM signals are the category most specific to Vasl’s target population and most absent from generic NLP training data. They represent culturally conditioned patterns of downplaying distress before disclosing it — a documented phenomenon in Black, Latino, and LGBTQIA+ communities where mental health help-seeking carries stigma. Generic models classify CCM signals as low risk. VLAP treats them as pre-disclosure markers that warrant gentle engagement prompts and session logging.

CCM-01VariableClassic minimization opener (“it's probably not a big deal but...”)
CCM-02VariableAAVE minimization constructions (“lowkey been struggling lately”)
CCM-03LowPermission-seeking before disclosure (“can I tell you something personal?”)
CCM-04ModerateGallows humor as disclosure mechanism (“not to be unalive about it but lmao”)
CCM-05ModerateAsking for a friend construction (“my friend is going through something...”)
CCM-06ModerateTopic drop without resolution (“actually forget I said anything”)
CCM-07LowInternalized stigma pattern (“other people have it so much worse”)
CCM-08LowPerformed resilience (“I'm strong, I've been through worse”)
Signal Example
"it's not that deep but lowkey been struggling."
CCM-01 + CCM-02 — Minimization opener predicts significant disclosure in what follows.
"God got me. That's all I'm saying."
PFE-01 territory if faith is being lost — here, CCM spiritual framing as conversational closure.
ISO-01 – ISO-07
7
Social Isolation & Withdrawal

Social isolation is a documented independent risk factor for suicidal ideation and is one of the two primary components of Joiner’s Interpersonal Theory of Suicide (thwarted belongingness). ISO signals are most clinically significant in combination with HOP or SHA signals, and as longitudinal markers of increasing withdrawal across sessions.

ISO-01ModerateDirect withdrawal statement (“not really talking to anyone these days”)
ISO-02HighRelationship severance language (“blocked a bunch of people”)
ISO-03ModeratePerceived burdensomeness, social expression (“I don't want to be a burden”)
ISO-04ModerateCode-switching exhaustion (“always performing, never actually me”)
ISO-05ModerateCommunity disconnection (“even my own people don't get it”)
ISO-06HighFamily estrangement, LGBTQIA+-specific (“got kicked out because of who I am”)
ISO-07LowDigital withdrawal references (“turned all my notifications off”)
Signal Example
"I'm moving different now. People can't handle it."
ISO withdrawal reframed as self-elevation — often follows rupture or rejection.
"My family doesn't know and if they found out..."
ISO-06 — Family estrangement, LGBTQIA+-specific. VLAP's highest-priority isolation signal.
TRM-01 – TRM-08
8
Trauma & Acute Stressor Markers

TRM signals capture social determinants of mental health that are disproportionately present in Vasl’s target population. Most TRM signals are not independently actionable clinical risk indicators — they are contextual elevators that increase the clinical significance of co-occurring HOP, SHA, or ISO signals. They are also the signals most normalized in community language and therefore most frequently missed by clinicians and tools not calibrated for this population.

TRM-01VariableCommunity violence exposure (“another one of my friends got shot”)
TRM-02ModerateHousing instability (“we got evicted so...”)
TRM-03VariableImmigration and documentation stress (“we don't know if we can stay”)
TRM-04LowFood insecurity references (“hadn't eaten since yesterday”)
TRM-05VariablePolice and legal system contact (“my brother just got locked up”)
TRM-06ModerateSchool discipline and push-out (“they're trying to expel me”)
TRM-07ModerateVicarious and cumulative loss (“I've been to too many funerals”)
TRM-08ModerateCaregiver incapacity / parentification (“I'm basically the parent right now”)
Signal Example
"Another one of my friends got shot."
TRM-01 — Community violence exposure. Not independently actionable; elevates co-occurring HOP/SHA signals.
"They're trying to expel me."
TRM-06 — School discipline and push-out. Documented trauma and mental health risk factor.
PFE-01 – PFE-07
7
Protective Factor Erosion

PFE signals are not distress signals in isolation — they indicate that a previously present protective factor has been removed, which elevates baseline risk. They are most clinically significant as longitudinal markers (V2 cross-session modeling) and in combination with HOP or SHA signals. The protective factors most relevant to Vasl’s population — faith community, mentors, peer networks, cultural identity, and aspirational goals — are different from those in the clinical literature developed for White, middle-class youth populations.

PFE-01LowLoss of faith or spiritual anchor (“don't really pray anymore”)
PFE-02ModerateTrusted adult loss (“the one adult who actually got me is gone”)
PFE-03ModeratePeer network dissolution (“everyone went their separate ways”)
PFE-04ModerateAchievement identity loss (“stopped performing, don't see the point”)
PFE-05LowFuture goal erasure (“those dreams don't seem realistic anymore”)
PFE-06LowCultural disconnection (“feel like I'm losing my culture”)
PFE-07LowLoss of routine and structure (“nothing feels normal or regular”)
Signal Example
"The one adult who actually got me is gone."
PFE-02 — Trusted adult loss. Removes a disproportionate share of the member's social safety net.
"Stopped thinking about that kind of future."
PFE-05 — Future goal erasure. A longitudinal hopelessness marker VLAP tracks across sessions.
Chapter 04 — What It Catches

Four patterns
standard AI
consistently misses.

These aren't edge cases. They're the language of the communities VLAP was built to serve. Standard sentiment analysis fails on all four — not because of a tuning problem, but because the training data never contained them. Each example below is a composite drawn from real signal categories in the taxonomy.

AAVE Hopelessness
HOP-03CCM-02
Member utterance — composite
"idk why i even try anymore tbh, can't keep doing this no more fr."
Indirect hopelessness expressed through AAVE constructions. "No more fr" (for real) is an authenticity escalator — not informal emphasis. Standard models penalize AAVE register as low-confidence text and miss the signal entirely.
Standard NLP: low-confidence / informal register — no distress signal generated.
Signal weight
High
Coded Self-Harm Language
SHA-02
Member utterance — composite
"been thinking about unaliving lately ngl."
"Unaliving" is a youth-community euphemism for self-harm ideation — developed specifically to avoid content moderation filters. It is in the VLAP vocabulary. Standard NLP models have never encountered it. The token does not exist in their training data.
Standard NLP: unknown token — utterance classified as ambiguous or neutral.
Signal weight
Critical
Pre-Disclosure Minimization
CCM-04ISO-04
Member utterance — composite
"it's not that deep but lowkey been struggling since school started."
CCM-04 reads "it's not that deep but" as a pre-disclosure minimization frame — a documented pattern in youth language where the speaker hedges before authentic disclosure. Standard models reduce signal weight at the minimization phrase. VLAP increases it: the hedge is the signal.
Standard NLP: minimization reduces overall distress classification. Marked as low-concern.
Signal weight
Notable
Burden Narrative
HOP-07
Member utterance — composite
"everything is fine now actually, I'm so over it lol."
Sudden positive shift following sustained distress signal history is classified as HOP-07 — a documented farewell and disengagement pattern. "Lol" here functions as a distancing marker, not genuine levity. Standard models classify this as sentiment improvement and deprioritize the member. VLAP treats it as an escalation signal.
Standard NLP: positive sentiment. Prior distress signals deprioritized. Status improved.
Signal weight
High
Chapter 05 — The Pipeline

From member
language to clinical
insight — in seven
accountable steps.

Every piece of text processed by VLAP moves through a documented, auditable pipeline. Each step has a clear technical purpose, a privacy control, and a human accountability point. No step produces a clinical decision. The pipeline produces context — and then a licensed clinician decides what to do with it.

Step 01 — Ingestion
Member language enters the system.

Text is received from Member App check-ins and coach messaging — channels where youth have explicitly chosen to share within a care relationship. No passive collection. No ambient monitoring. Language enters the pipeline only through intentional member action within the Vasl platform.

Channel: Member App · Coach messaging · Consent-gated input only
Active input only
Step 02 — Consent Gate
Consent is verified before any processing occurs.

Before any inference runs, the system verifies that the member has valid, current consent for language processing. No consent — the text is rejected immediately and never enters the inference pipeline. Consent is not assumed from enrollment. It is verified at the point of processing, for every submission.

Consent record checked against member profile · No consent = immediate rejection · Audit logged
Hard gate
Step 03 — PII Removal
All identifying information is stripped before the model sees a single word.

Microsoft Presidio — an enterprise-grade PII detection and removal system — scrubs all personally identifiable information from the text before it reaches the inference layer. Names, locations, phone numbers, account references, and identifying context are removed. The model never processes a member's identity — only their language.

Microsoft Presidio · Named entity recognition · Regex + ML detection · Output verified before passing
Identity-blind inference
Step 04 — Tokenization & Vocabulary Matching
The extended vocabulary layer activates — reading language that standard models cannot.

The text is tokenized using VLAP's extended vocabulary — which includes the 2,400+ AAVE and youth vernacular tokens added to the base vocabulary. This is where standard NLP fails: it encounters tokens like "unaliving" or constructions like "can't keep doing this no more fr" and has no representation for them. VLAP has trained representations for all of them.

Extended BERT-family tokenizer · 2,400+ custom tokens · Cultural register detection active
Cultural vocabulary
Step 05 — Inference
VLAP classifies signals across all clinical categories.

The model runs inference across the full VLAP Distress Signal Taxonomy — evaluating the input against all 47 signals across six clinical categories simultaneously. CCM modifiers adjust signal weights based on cultural register, minimization patterns, and code-switching context. Output is a structured dimensional signal profile — not a score, not a diagnosis, not a risk tier.

VLAP V1 · CCM weighting active
Non-diagnostic output
Step 06 — Clinical Output Generation
Structured interpretive context is formatted for the clinician — and only the clinician.

The dimensional signal profile is structured into an interpretive context package — signal codes, community-language annotations, confidence indicators, and session history context — and delivered to the clinician's Coach Portal view. The raw text is discarded at this point. What the clinician receives is context, not content. The member's words do not persist in the system.

Coach Portal delivery · Raw text discarded · Signal profile retained for 90 days · Audit log: 6 years
Clinician-only delivery
Step 07 — Human Review
A licensed clinician receives the context and determines the response.

The pipeline ends here — with a human being. No automated action follows from VLAP output. No alert fires. No protocol activates. The clinician reviews the dimensional signal context before the session and decides what to do with it. For signals or signal combinations that meet High-risk escalation criteria, a clinical supervisor review is initiated within 90 minutes — by a person, not a trigger.

Licensed clinician review required · High-risk escalation: clinical supervisor + 90-min human SLA · No automated action
Human determination
Chapter 06 — Design Principles

Four principles
that cannot be
traded away.

These are not guidelines. They are the non-negotiable constraints that governed every architectural decision in VLAP — from training data selection through deployment. If a design choice conflicted with any of these, the design changed. Not the principle.

01
Never member-facing. Not now, not ever.

No member ever sees a signal code, a confidence weight, or any output from VLAP. Not a simplified version. Not a wellness score derived from it. Nothing. VLAP outputs flow exclusively to licensed clinicians and authorized clinical administrators. The member experiences the care that follows — not the system behind it.

Technical Implementation
Coach Portal access requires licensed clinician credential verification. Member-facing app has zero API surface to VLAP output. No derived scores, no wellness indicators, no secondary displays of VLAP inference. Access control is architectural, not policy-based — member app and clinical portal are separate systems.
02
Human in the loop. Always. No exceptions.

No automated action follows from VLAP output. A clinical supervisor does not receive an automated directive. The 988 crisis line is not called by the platform. A coach is not dispatched by an algorithm. What VLAP does is inform a human — who then decides what to do. The human is not a checkpoint in an automated system. The human is the system.

Technical Implementation
All VLAP output is surfaced to the clinical dashboard as read-only context. No output field connects to any automated action, notification dispatch, or external API call. High-risk escalation signals generate a review queue item — not an alert or an automated escalation. The queue requires a licensed clinician to open, review, and document a response.
03
Zero raw text storage. The member's words don't live in the system.

Member language is processed in-memory through the VLAP pipeline and discarded at the point of clinical output generation. What persists is the dimensional signal profile — not the text that generated it. The member's words are not in a database. The meaning the model extracted from them is, in a de-identified form that serves only the care relationship it was created in.

Technical Implementation
In-memory processing pipeline. Raw text buffer cleared at Step 06 output generation. Signal profile retention: 90 days, tied to member care record. Audit log retention: 6 years per HIPAA requirement. No raw text in backup systems, model retraining pipelines, or analytics infrastructure.
04
Bias monitoring is not a metric. It's an operational gate.

False positive and false negative rates are tracked by demographic subgroup — race, ethnicity, gender identity, age cohort — in every inference batch. If systematic disparities emerge, they don't go into a report. They trigger a mandatory model review that suspends deployment of the affected signal until the disparity is resolved. This is how a model built for equity stays that way.

Technical Implementation
Automated demographic disaggregation of FPR/FNR in every inference batch. Disparity threshold: >5% divergence across subgroups triggers mandatory review. Review team: clinical staff + community advisor panel + model engineering. Affected signal suspended from clinical output during review. Finding documented and published in bias monitoring log accessible to all clinical partners.
Chapter 07 — Validation & Compliance

In active
clinical validation.
Not just
deployed.

VLAP is not waiting for validation to come back from a lab. It is in active clinical deployment — with a parallel IRB study underway at an independent academic research institution validating signal accuracy in production conditions. The model earns its claim in the real world, with real youth, under real clinical oversight. Results will be published upon study completion.

Active Study — IRB Approved
Clinical Signal Accuracy Validation — VASL-IRB-001
IRB-approved study validating VLAP signal accuracy against licensed clinician assessment in production deployment conditions. Study evaluates signal sensitivity, specificity, and demographic subgroup parity across Black, Latino, LGBTQ+, and youth in urban and rural communities.
Study in progress — results to be published upon completion
HIPAA
Full technical safeguard implementation. BAA required for all partners.
Complete HIPAA technical safeguard mapping across all platform components. Business Associate Agreement required for every clinical partner deployment. Annual compliance review. Audit log retention: 6 years.
SOC 2 Type II
Annual third-party security audit. Full report available under NDA.
SOC 2 Type II certification covering security, availability, and confidentiality trust service criteria. Third-party audit conducted annually. Report available to clinical partners under NDA upon request.
FERPA
Student records handled per FERPA requirements for school-based deployments.
School-based program deployments operate under FERPA-compliant data handling protocols. Student health records segregated from academic records. Consent architecture designed for minor students and their guardians.
Not a Medical Device
Clinical decision support software. All outputs require clinician review before any action.
VLAP is classified as clinical decision support software — not a medical device under current FDA guidance. All model outputs require licensed clinician review and determination before any clinical action is taken. The model does not diagnose, prescribe, or direct care.
Chapter 08 — Technical Specification

37 pages.
Every architectural
decision documented.

The VLAP V1 Technical Specification covers model architecture, training data methodology, API contract, signal taxonomy definitions, bias monitoring protocols, and deployment acceptance criteria. It exists because any clinical partner deploying this platform deserves to understand exactly what they're deploying.

§01
Model Architecture & Base Training
Pages 1–6 · BERT-family architecture · Fine-tuning methodology
§02
Training Corpus & Annotation Methodology
Pages 7–12 · Inter-annotator agreement · Community review process
§03
VLAP Distress Signal Taxonomy v1.0 — Full Definitions
Pages 13–20 · 47 signal definitions · Escalation rules · Confidence weighting
§04
API Contract & Integration Requirements
Pages 21–26 · Endpoint documentation · Auth · Rate limits · Response schema
§05
Bias Monitoring Protocols & Disparity Thresholds
Pages 27–32 · Demographic disaggregation · FPR/FNR tracking · Review triggers
§06
Deployment Acceptance Criteria & Clinical Onboarding
Pages 33–37 · Partner requirements · Clinical staff training · Go-live criteria
Related Reading
the bias problem this model addresses
Differential item functioning, dialect bias, and why standard screening instruments don't hold up across groups.
VLAP, the layer this model powers
How VLAP's output becomes the signal taxonomy a human actually reviews.
the architecture this model runs inside
PHI boundaries and the security model around every prediction this model makes.