Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Salience and Reward Prediction Error

Domaine:

natural language processing
Créateur:
Ben
Éditeur:
Cen
Éditeur:
OSF
Hôte:avatar
NOTE: entries are written in past tense. This is purely for convenience, as much of the text will be reused in the final publication, where past tense is necessary. This preregistration was completed before any data was collected. Two separate word sets were created for this experiment. The first set contained 1296 Swahili words taken from kamusi.org. All words were two syllables long. The second set contained 324 English words and were taken from the SUBTLEX US word frequency database (CITATION). All words were four syllables long, and had a zipf value between 3.3 and 3.6, indicating medium frequency of occurrence (CITATION). This meant that words were matched not only for length but also for how common they are. Four syllable English words were chosen to make memorization of word-pairs more difficult, so that the recognition task would remain difficult despite giving only two options per trial. Medium word frequency was chosen to help make the task more difficult while also having words be reasonably recognizable, unlike low frequency words which could be completely unknown to the participant. The two word sets were created completely independently, meaning that words were chosen without considering whether there is a corresponding translation in the other word set. In all likelihood there are no true English translations of Swahili words contained within the pair of word sets. Acquisition task. This task was taken from de Loof et al. (2018). At the start of each trial, the screen displayed one English word and four Swahili words. The words stayed on screen for four seconds, after which a frame appeared around the possible Swahili translations for the English word. Either one or four words were framed. The framed words indicated possible correct translations. In the one-option condition, participants were guaranteed to answer correctly. In the four-option condition participants had a 25% chance of choosing the correct translation. Each word position was assigned to one keyboard button (‘f’, ‘v’, ‘j’, or ‘n’). Participants were allowed to take as long as they wanted when deciding. Participants were told to answer based on their first impression. Participants were then given feedback to indicate whether they had answered correctly. However, the feedback was secretly preprogrammed, so that a fixed number of trials would be rewarded. This meant that participants were not necessarily learning real translations of English words. For example, if a trial had been preprogrammed to have four options and be rewarded, then participants would receive positive feedback for their answer, regardless of which Swahili word they selected. This chosen word would become the word they had to memorize. This made sure we had enough of every option-reward combination (rewarded with one option, rewarded and unrewarded with four options). Furthermore, for each English word, the four Swahili words were randomly selected, and likely did not include the actual Swahili translation. This ensured that any linguistic regularity between the English and Swahili words could not affect participants’ learning. The feedback given to participants included the English word, an equal sign and the (ostensibly) correct Swahili word displayed in the middle of the screen. During rewarded trials, a green frame appeared around the English word and the chosen Swahili word. During unrewarded trials, a red frame appeared around the English word and one of the unselected Swahili words. The words stayed on screen for five seconds. Participants were told to use this time to memorize the word pair being shown. In total, 108 trials were completed. Every Swahili word chosen by a participant was given an RPE. Our calculation of RPE follows that of De Loof et al. (2018). The obtained reward of a given trial was 1 on rewarded trials and 0 on unrewarded trials. The reward probability of a given trial was 1 when one possible translation was shown, and 0.25 when four possible translations were shown. RPE was calculated by taking the obtained reward and subtracting the reward probability. As such, each combination of reward and number of options gave a unique RPE ranging from -0.25 to 0.75. As noted during the introduction, there is some debate whether absolute values of RPEs should be used instead of actual values. This question is beyond the scope of our study. Filler Task Before beginning the recognition test, participants completed a filler task to reduce recency effects. As in de Loof et al. (2018), a magnitude comparison task was used, in which participants categorized 400 numbers between 1 and 9 (not including 5). Participants pressed the ‘f’ button when the number was smaller than 5, and the ‘j’ button when the number was larger than 5. Recognition test Before the recognition test, participants were warned that some trials would not have a correct answer, though most trials would. For each trial, an English word appeared at the top of the screen along with two Swahili words below. As soon as the words appeared, participants could choose between the two Swahili words by using ‘f’ and ‘j’ to select the left and right word, respectively. No time constraints were given. Each trial presented one distractor word and one target word. The distractor word was always a word that had been previously rewarded during a one-option trial (meaning it had an RPE of 0). In total, there are three possible target words. In trial type A, the target word had an RPE of 0 and was the correct translation. In trial type B, the target word had an RPE of 0.75 (meaning that it was a rewarded answer among four options during the learning phase). It was also the correct translation. In trial type C, the target word had an RPE of 0.75. However, it was not the correct translation. Therefore, in trial type C, there was no correct translation available. Note that the trial type labels are purely to help describe the task. In practice the participant was not told about the concept of trial types, and the order that trials were presented was randomized. In total, 27 trials were presented (9 of each trial type). No feedback was given after recognition trials. After each trial, participants were shown a Likert scale asking how confident they were in their answer. The options included “very unconfident”, “somewhat unconfident”, “somewhat confident”, “very confident”. Participants answered by pressing a number from 1 to 4. Repeat of Experiment Unfortunately, a single session of data collection only gives 27 recognition trials. This is not enough to estimate participant-level random effects, which we needed for data analysis. We could not simply increase the number of trials per session, as that risked burnout, where participants would stop learning as effectively due to mental exhaustion. Therefore, we decided to have each participant complete three days of testing. Each day involved the same experiment, just with different words. The words for each day were randomly chosen during day one of each experiment. This meant that each participant would get a different set of words per day, so there was no connection between experiment day and words shown. Each word was used exactly once, so there was no repeating words and no unused words. In this way, each participant was shown every word in a random order. We therefore chose to combine data from all three days into a single overall dataset per participant. This allowed us to get 81 data points per participant, which is sufficient for the analysis models we used. No random effects were included to distinguish what day data was collected on, as this would force our model to calculate an effect for each participant/day combination separately, which leads back to the issue of estimating random effects with only 27 trials. This does mean that our study cannot make claims about when the effect of RPE takes place. However, we do not feel that this is relevant to our research question. If an effect of RPE is found, but (unbeknownst to us) it only exists in the second and third day, this does not make conclusions about the effect of RPE invalid. Rather, it means that future study are needed to determine how the effect of RPE interacts with time.

Visit

doi.orgosf.io

Languages

Swahili

Tags

Social and Behavioral SciencesConjoint AnalysisDeclarative learningPsychologyFOS: PsychologyReward Prediction ErrorsSalience

Licenses

No license