Logo Lanfrica

sirenism/vanda-dataset

Domaine:

natural language processing

Type de record:

dataset
Créateur:
sir
Hôte:
# Vanda Dataset This is the official repository of **VANDAlize: Safety Degradation of Large Language Models Through Taglish-Based Inputs**. ## Keywords `LLM`, `Code-switching`, `Multilinggual`, `Red-teaming` ## Abstract tbd ## Data Sources The English translations of the harmful prompts subset and their respective harm categories were taken from the Aya Red-teaming Dataset. The English and Filipino translations of the benign prompts subset were taken from the Tagalog-Filipino-English-Translation Dataset.

Languages

Licenses