Logo Lanfrica

sirenism/vanda-dataset

Domain:

natural language processing

Record type:

dataset
Creator:
sir
Host:
# Vanda Dataset This is the official repository of **VANDAlize: Safety Degradation of Large Language Models Through Taglish-Based Inputs**. ## Keywords `LLM`, `Code-switching`, `Multilinggual`, `Red-teaming` ## Abstract tbd ## Data Sources The English translations of the harmful prompts subset and their respective harm categories were taken from the Aya Red-teaming Dataset. The English and Filipino translations of the benign prompts subset were taken from the Tagalog-Filipino-English-Translation Dataset.

Languages

Licenses