VEILED: Vulnerability expression through indirect language evaluation dataset
arXiv, 2026
Tan et al. (2026). "VEILED: Vulnerability expression through indirect language evaluation dataset." arXiv. https://openreview.net/pdf?id=VsLiSIS5VE
AI systems increasingly encounter users who express distress indirectly, through culturally situated language, euphemism, or allusion. They may miss these signals or provide support mismatched to the application context. Existing benchmarks rarely expose such failures: standardized prompts flatten cultural variation, and single ground truths obscure practitioner disagreement. We present VEILED (Vulnerability Expression through Indirect Language Evaluation Dataset), a benchmark and construction methodology for evaluating single-turn responses to culturally situated expressions of suicide, self-harm, and distress as proxy alignment with calibrated practitioner references. VEILED provides prompts across four contexts (counselor chat, forum post, platform search, and triage tool), an expert-validated multi-label taxonomy separating distress intent, communicative function, and response behavior, and a plural-reference evaluation method that preserves practitioner disagreement. In an empirical demonstration on five frontier chat systems, changing the practitioner reference shifts absolute proxy policy alignment by 20–40%, while the relative ordering of practitioner references by alignment is consistent across systems.